Carlos Gómez-Rodríguez

dblp:95/3319 · DBLP profile ↗
← Back
65ranked-venue papers
20as first author
15since 2021 · last 2026
0000-0003-0752-8812ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 57 · 17 first-author · 15 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Theory of computation · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 More Aligned, Less Diverse? Analyzing the Grammar and Lexicon of Two Generations of LLMs
abstract
Adrián Gude, Roi Santos-Rios, Francis Bond, Dan Flickinger, Carlos Gómez-Rodríguez, Olga Zamaraeva. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Adrián Gude, Roi Santos-Rios, Francis Bond, Dan Flickinger, Carlos Gómez-Rodríguez, Olga Zamaraeva
ACL (1)5
2026 HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning
Alexis Correa-Guillén, Carlos Gómez-Rodríguez, David Vilares 0001
LREC2
2025 Hierarchical Bracketing Encodings for Dependency Parsing as Tagging
abstract
We present a family of encodings for sequence labeling dependency parsing, based on the concept of hierarchical bracketing. We prove that the existing 4-bit projective encoding belongs to this family, but it is suboptimal in the number of labels used to encode a tree. We derive an optimal hierarchical bracketing, which minimizes the number of symbols used and encodes projective trees using only 12 distinct labels (vs. 16 for the 4-bit encoding). We also extend optimal hierarchical bracketing to support arbitrary non-projectivity in a more compact way than previous encodings. Our new encodings yield competitive accuracy on a diverse set of treebanks.
Ana Ezquerro, David Vilares 0001, Anssi Yli-Jyrä, Carlos Gómez-Rodríguez
ACL (1)4
2025 Comparing LLM-generated and human-authored news text using formal syntactic theory
abstract
This study provides the first comprehensive comparison of New York Times-style text generated by six large language models against real, human-authored NYT writing. The comparison is based on a formal syntactic theory. We use Head-driven Phrase Structure Grammar (HPSG) to analyze the grammatical structure of the texts. We then investigate and illustrate the differences in the distributions of HPSG grammar types, revealing systematic distinctions between human and LLM-generated writing. These findings contribute to a deeper understanding of the syntactic behavior of LLMs as well as humans, within the NYT genre.
Olga Zamaraeva, Dan Flickinger, Francis Bond, Carlos Gómez-Rodríguez
ACL (1)4
2025 Hierarchical Bracketing Encodings Work for Dependency Graphs
abstract
We revisit hierarchical bracketing encodings from a practical perspective in the context of dependency graph parsing.The approach encodes graphs as sequences, enabling lineartime parsing with n tagging actions, and still representing reentrancies, cycles, and empty nodes.Compared to existing graph linearizations, this representation substantially reduces the label space while preserving structural information.We evaluate it on a multilingual and multi-formalism benchmark, showing competitive results and consistent improvements over other methods in exact match accuracy.
Ana Ezquerro, Carlos Gómez-Rodríguez, David Vilares 0001
EMNLP2
2024 Spanish Resource Grammar Version 2023
abstract
We present the latest version of the Spanish Resource Grammar (SRG), a grammar of Spanish implemented in the HPSG formalism. Such grammars encode a complex set of hypotheses about syntax making them a resource for empirical testing of linguistic theory. They also encode a strict notion of grammaticality which makes them a resource for natural language processing applications in computer-assisted language learning. This version of the SRG uses the recent version of the Freeling morphological analyzer and is released along with an automatically created, manually verified treebank of 2,291 sentences. We explain the treebanking process, emphasizing how it is different from treebanking with manual annotation and how it contributes to empirically-driven development of syntactic theory. The treebanks’ high level of consistency and detail makes them a resource for training high-quality semantic parsers and generally systems that benefit from precise and detailed semantics. Finally, we present the grammar’s coverage and overgeneration on 100 sentences from a learner corpus, a new research line related to developing methodologies for robust empirical evaluation of hypotheses in second language acquisition.
Olga Zamaraeva, Lorena S. Allegue, Carlos Gómez-Rodríguez
LREC/COLING3
2024 Dependency Graph Parsing as Sequence Labeling
abstract
Various linearizations have been proposed to cast syntactic dependency parsing as sequence labeling.However, these approaches do not support more complex graph-based representations, such as semantic dependencies or enhanced universal dependencies, as they cannot handle reentrancy or cycles.By extending them, we define a range of unbounded and bounded linearizations that can be used to cast graph parsing as a tagging task, enlarging the toolbox of problems that can be solved under this paradigm.Experimental results on semantic dependency and enhanced UD parsing show that with a good choice of encoding, sequencelabeling dependency graph parsers combine high efficiency with accuracies close to the state of the art, in spite of their simplicity.
Ana Ezquerro, David Vilares 0001, Carlos Gómez-Rodríguez
EMNLP3
2024 Revisiting Supertagging for faster HPSG parsing
abstract
We present new supertaggers trained on English grammar-based treebanks and test the effects of the best tagger on parsing speed and accuracy.The treebanks are produced automatically by large manually built grammars and feature high-quality annotation based on a well-developed linguistic theory (HPSG).The English Resource Grammar treebanks include diverse and challenging test datasets, beyond the usual WSJ section 23 and Wikipedia data.HPSG supertagging has previously relied on MaxEnt-based models.We use SVM and neural CRF-and BERT-based methods and show that both SVM and neural supertaggers achieve considerably higher accuracy compared to the baseline and lead to an increase not only in the parsing speed but also the parser accuracy with respect to gold dependency structures.Our fine-tuned BERT-based tagger achieves 97.26% accuracy on 950 sentences from WSJ23 and 93.88% on the out-of-domain technical essay The Cathedral and the Bazaar (cb)).We present experiments with integrating the best supertagger into an HPSG parser and observe a speedup of a factor of 3 with respect to the system which uses no tagging at all, as well as large recall gains and an overall precision gain.We also compare our system to an existing integrated tagger and show that although the well-integrated tagger remains the fastest, our experimental system can be more accurate.Finally, we hope that the diverse and difficult datasets we used for evaluation will gain more popularity in the field: we show that results can differ depending on the dataset, even if it is an in-domain one.We contribute the complete datasets reformatted for Huggingface token classification.
Olga Zamaraeva, Carlos Gómez-Rodríguez
EMNLP2
2023 4 and 7-bit Labeling for Projective and Non-Projective Dependency Trees
abstract
We introduce an encoding for syntactic parsing as sequence labeling that can represent any projective dependency tree as a sequence of 4-bit labels, one per word.The bits in each word's label represent (1) whether it is a right or left dependent, (2) whether it is the outermost (left/right) dependent of its parent, (3) whether it has any left children and ( 4) whether it has any right children.We show that this provides an injective mapping from trees to labels that can be encoded and decoded in linear time.We then define a 7-bit extension that represents an extra plane of arcs, extending the coverage to almost full non-projectivity (over 99.9% empirical arc coverage).Results on a set of diverse treebanks show that our 7-bit encoding obtains substantial accuracy gains over the previously best-performing sequence labeling encodings.
Carlos Gómez-Rodríguez, Diego Roca, David Vilares 0001
EMNLP1
2023 Assessment of Pre-Trained Models Across Languages and Grammars
abstract
Alberto Muñoz-Ortiz, David Vilares, Carlos Gómez-Rodríguez. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Alberto Muñoz-Ortiz, David Vilares 0001, Carlos Gómez-Rodríguez
IJCNLP (1)3
2023 Discontinuous grammar as a foreign language
abstract
In order to achieve deep natural language understanding, syntactic constituent parsing is a vital step, highly demanded by many artificial intelligence systems to process both text and speech. One of the most recent proposals is the use of standard sequence-to-sequence models to perform constituent parsing as a machine translation task, instead of applying task-specific parsers. While they show a competitive performance, these text-to-parse transducers are still lagging behind classic techniques in terms of accuracy, coverage and speed. To close the gap, we here extend the framework of sequence-to-sequence models for constituent parsing, not only by providing a more powerful neural architecture for improving their performance, but also by enlarging their coverage to handle the most complex syntactic phenomena: discontinuous structures. To that end, we design several novel linearizations that can fully produce discontinuities and, for the first time, we test a sequence-to-sequence model on the main discontinuous benchmarks, obtaining competitive results on par with task-specific discontinuous constituent parsers and achieving state-of-the-art scores on the (discontinuous) English Penn Treebank.
Daniel Fernández-González, Carlos Gómez-Rodríguez
Neurocomputing2
2022 The Fragility of Multi-Treebank Parsing Evaluation
abstract
Treebank selection for parsing evaluation and the spurious effects that might arise from a biased choice have not been explored in detail. This paper studies how evaluating on a single subset of treebanks can lead to weak conclusions. First, we take a few contrasting parsers, and run them on subsets of treebanks proposed in previous work, whose use was justified (or not) on criteria such as typology or data scarcity. Second, we run a large-scale version of this experiment, create vast amounts of random subsets of treebanks, and compare on them many parsers whose scores are available. The results show substantial variability across subsets and that although establishing guidelines for good treebank selection is hard, some inadequate strategies can be easily avoided.
Iago Alonso-Alonso, David Vilares 0001, Carlos Gómez-Rodríguez
COLING3
2022 The Impact of Edge Displacement Vaserstein Distance on UD Parsing Performance
abstract
Abstract We contribute to the discussion on parsing performance in NLP by introducing a measurement that evaluates the differences between the distributions of edge displacement (the directed distance of edges) seen in training and test data. We hypothesize that this measurement will be related to differences observed in parsing performance across treebanks. We motivate this by building upon previous work and then attempt to falsify this hypothesis by using a number of statistical methods. We establish that there is a statistical correlation between this measurement and parsing performance even when controlling for potential covariants. We then use this to establish a sampling technique that gives us an adversarial and complementary split. This gives an idea of the lower and upper bounds of parsing systems for a given treebank in lieu of freshly sampled data. In a broader sense, the methodology presented here can act as a reference for future correlation-based exploratory work in NLP.
Mark Anderson 0005, Carlos Gómez-Rodríguez
Comput. Linguistics2
2022 Multitask Pointer Network for multi-representational parsing
abstract
Dependency and constituent trees are widely used by many artificial intelligence applications for representing the syntactic structure of human languages. Typically, these structures are separately produced by either dependency or constituent parsers. In this article, we propose a transition-based approach that, by training a single model, can efficiently parse any input sentence with both constituent and dependency trees, supporting both continuous/projective and discontinuous/non-projective syntactic structures. To that end, we develop a Pointer Network architecture with two separate task-specific decoders and a common encoder, and follow a multitask learning strategy to jointly train them. The resulting quadratic system, not only becomes the first parser that can jointly produce both unrestricted constituent and dependency trees from a single model, but also proves that both syntactic formalisms can benefit from each other during training, achieving state-of-the-art accuracies in several widely-used benchmarks such as the continuous English and Chinese Penn Treebanks, as well as the discontinuous German NEGRA and TIGER datasets.
Daniel Fernández-González, Carlos Gómez-Rodríguez
Knowl. Based Syst.2
2021 Reducing Discontinuous to Continuous Parsing with Pointer Network Reordering
abstract
Discontinuous constituent parsers have always lagged behind continuous approaches in terms of accuracy and speed, as the presence of constituents with discontinuous yield introduces extra complexity to the task.However, a discontinuous tree can be converted into a continuous variant by reordering tokens.Based on that, we propose to reduce discontinuous parsing to a continuous problem, which can then be directly solved by any off-the-shelf continuous parser.To that end, we develop a Pointer Network capable of accurately generating the continuous token arrangement for a given input sentence and define a bijective function to recover the original order.Experiments on the main benchmarks with two continuous parsers prove that our approach is on par in accuracy with purely discontinuous state-of-the-art algorithms, but considerably faster.
Daniel Fernández-González, Carlos Gómez-Rodríguez
EMNLP (1)2
2020 Discontinuous Constituent Parsing with Pointer Networks
abstract
One of the most complex syntactic representations used in computational linguistics and NLP are discontinuous constituent trees, crucial for representing all grammatical phenomena of languages such as German. Recent advances in dependency parsing have shown that Pointer Networks excel in efficiently parsing syntactic relations between words in a sentence. This kind of sequence-to-sequence models achieve outstanding accuracies in building non-projective dependency trees, but its potential has not been proved yet on a more difficult task. We propose a novel neural network architecture that, by means of Pointer Networks, is able to generate the most accurate discontinuous constituent representations to date, even without the need of Part-of-Speech tagging information. To do so, we internally model discontinuous constituent structures as augmented non-projective dependency structures. The proposed approach achieves state-of-the-art results on the two widely-used NEGRA and TIGER benchmarks, outperforming previous work by a wide margin.
Daniel Fernández-González, Carlos Gómez-Rodríguez
AAAI2
2020 Parsing as Pretraining
abstract
Recent analyses suggest that encoders pretrained for language modeling capture certain morpho-syntactic structure. However, probing frameworks for word vectors still do not report results on standard setups such as constituent and dependency parsing. This paper addresses this problem and does full parsing (on English) relying only on pretraining architectures – and no decoding. We first cast constituent and dependency parsing as sequence tagging. We then use a single feed-forward layer to directly map word vectors to labels that encode a linearized tree. This is used to: (i) see how far we can reach on syntax modelling with just pretrained encoders, and (ii) shed some light about the syntax-sensitivity of different word vectors (by freezing the weights of the pretraining network during training). For evaluation, we use bracketing F1-score and las, and analyze in-depth differences across representations for span lengths and dependency displacements. The overall results surpass existing sequence tagging parsers on the ptb (93.5%) and end-to-end en-ewt ud (78.8%).
David Vilares 0001, Michalina Strzyz, Anders Søgaard, Carlos Gómez-Rodríguez
AAAI4
2020 Enriched In-Order Linearization for Faster Sequence-to-Sequence Constituent Parsing
abstract
Sequence-to-sequence constituent parsing requires a linearization to represent trees as sequences.Top-down tree linearizations, which can be based on brackets or shift-reduce actions, have achieved the best accuracy to date.In this paper, we show that these results can be improved by using an inorder linearization instead.Based on this observation, we implement an enriched inorder shift-reduce linearization inspired by Vinyals et al. (2015)'s approach, achieving the best accuracy to date on the English PTB dataset among fully-supervised single-model sequence-to-sequence constituent parsers.Finally, we apply deterministic attention mechanisms to match the speed of state-of-theart transition-based parsers, thus showing that sequence-to-sequence models can match them, not only in accuracy, but also in speed.
Daniel Fernández-González, Carlos Gómez-Rodríguez
ACL2
2020 Transition-based Semantic Dependency Parsing with Pointer Networks
abstract
Transition-based parsers implemented with Pointer Networks have become the new state of the art in dependency parsing, excelling in producing labelled syntactic trees and outperforming graph-based models in this task.In order to further test the capabilities of these powerful neural networks on a harder NLP problem, we propose a transition system that, thanks to Pointer Networks, can straightforwardly produce labelled directed acyclic graphs and perform semantic dependency parsing.In addition, we enhance our approach with deep contextualized word embeddings extracted from BERT.The resulting system not only outperforms all existing transitionbased models, but also matches the best fullysupervised accuracy to date on the SemEval 2015 Task 18 English datasets among previous state-of-the-art graph-based parsers.
Daniel Fernández-González, Carlos Gómez-Rodríguez
ACL2
2020 Data Augmentation via Subtree Swapping for Dependency Parsing of Low-Resource Languages
abstract
The lack of annotated data is a big issue for building reliable NLP systems for most of the world's languages.But this problem can be alleviated by automatic data generation.In this paper, we present a new data augmentation method for artificially creating new dependency-annotated sentences.The main idea is to swap subtrees between annotated sentences while enforcing strong constraints on those trees to ensure maximal grammaticality of the new sentences.We also propose a method to perform low-resource experiments using resource-rich languages by mimicking low-resource languages by sampling sentences under a low-resource distribution.In a series of experiments, we show that our newly proposed data augmentation method outperforms previous proposals using the same basic inputs.
Mathieu Dehouck, Carlos Gómez-Rodríguez
COLING2
2020 A Unifying Theory of Transition-based and Sequence Labeling Parsing
abstract
We define a mapping from transition-based parsing algorithms that read sentences from left to right to sequence labeling encodings of syntactic trees.This not only establishes a theoretical relation between transition-based parsing and sequence-labeling parsing, but also provides a method to obtain new encodings for fast and simple sequence labeling parsing from the many existing transition-based parsers for different formalisms.Applying it to dependency parsing, we implement sequence labeling versions of four algorithms, showing that they are learnable and obtain comparable performance to existing encodings.
Carlos Gómez-Rodríguez, Michalina Strzyz, David Vilares 0001
COLING1
2020 Bracketing Encodings for 2-Planar Dependency Parsing
abstract
We present a bracketing-based encoding that can be used to represent any 2-planar dependency tree over a sentence of length n as a sequence of n labels, hence providing almost total coverage of crossing arcs in sequence labeling parsing.First, we show that existing bracketing encodings for parsing as labeling can only handle a very mild extension of projective trees.Second, we overcome this limitation by taking into account the well-known property of 2-planarity, which is present in the vast majority of dependency syntactic structures in treebanks, i.e., the arcs of a dependency tree can be split into two planes such that arcs in a given plane do not cross.We take advantage of this property to design a method that balances the brackets and that encodes the arcs belonging to each of those planes, allowing for almost unrestricted non-projectivity (∼ 99.9% coverage) in sequence labeling parsing.The experiments show that our linearizations improve over the accuracy of the original bracketing encoding in highly non-projective treebanks (on average by 0.4 LAS), while achieving a similar speed.Also, they are especially suitable when PoS tags are not used as input parameters to the models.
Michalina Strzyz, David Vilares 0001, Carlos Gómez-Rodríguez
COLING3
2020 On the Frailty of Universal POS Tags for Neural UD Parsers
abstract
We present an analysis on the effect UPOS accuracy has on parsing performance.Results suggest that leveraging UPOS tags as features for neural parsers requires a prohibitively high tagging accuracy and that the use of gold tags offers a non-linear increase in performance, suggesting some sort of exceptionality.We also investigate what aspects of predicted UPOS tags impact parsing accuracy the most, highlighting some potentially meaningful linguistic facets of the problem.
Mark Anderson 0005, Carlos Gómez-Rodríguez
CoNLL2
2020 Discontinuous Constituent Parsing as Sequence Labeling
abstract
This paper reduces discontinuous parsing to sequence labeling.It first shows that existing reductions for constituent parsing as labeling do not support discontinuities.Second, it fills this gap and proposes to encode tree discontinuities as nearly ordered permutations of the input sequence.Third, it studies whether such discontinuous representations are learnable.The experiments show that despite the architectural simplicity, under the right representation, the models are fast and accurate.1
David Vilares 0001, Carlos Gómez-Rodríguez
EMNLP (1)2
2020 Inherent Dependency Displacement Bias of Transition-Based Algorithms
abstract
A wide variety of transition-based algorithms are currently used for dependency parsers. Empirical studies have shown that performance varies across different treebanks in such a way that one algorithm outperforms another on one treebank and the reverse is true for a different treebank. There is often no discernible reason for what causes one algorithm to be more suitable for a certain treebank and less so for another. In this paper we shed some light on this by introducing the concept of an algorithm’s inherent dependency displacement distribution. This characterises the bias of the algorithm in terms of dependency displacement, which quantify both distance and direction of syntactic relations. We show that the similarity of an algorithm’s inherent distribution to a treebank’s displacement distribution is clearly correlated to the algorithm’s parsing performance on that treebank, specificially with highly significant and substantial correlations for the predominant sentence lengths in Universal Dependency treebanks. We also obtain results which show a more discrete analysis of dependency displacement does not result in any meaningful correlations.
Mark Anderson 0005, Carlos Gómez-Rodríguez
LREC2
2020 Cross-Lingual Word Embeddings for Turkic Languages
abstract
There has been an increasing interest in learning cross-lingual word embeddings to transfer knowledge obtained from a resource-rich language, such as English, to lower-resource languages for which annotated data is scarce, such as Turkish, Russian, and many others. In this paper, we present the first viability study of established techniques to align monolingual embedding spaces for Turkish, Uzbek, Azeri, Kazakh and Kyrgyz, members of the Turkic family which is heavily affected by the low-resource constraint. Those techniques are known to require little explicit supervision, mainly in the form of bilingual dictionaries, hence being easily adaptable to different domains, including low-resource ones. We obtain new bilingual dictionaries and new word embeddings for these languages and show the steps for obtaining cross-lingual word embeddings using state-of-the-art techniques. Then, we evaluate the results using the bilingual dictionary induction task. Our experiments confirm that the obtained bilingual dictionaries outperform previously-available ones, and that word embeddings from a low-resource language can benefit from resource-rich closely-related languages when they are aligned together. Furthermore, evaluation on an extrinsic task (Sentiment analysis on Uzbek) proves that monolingual word embeddings can, although slightly, benefit from cross-lingual alignments.
Elmurod Kuriyozov, Yerai Doval, Carlos Gómez-Rodríguez
LREC3
2019 Sequence Labeling Parsing by Learning across Representations
abstract
We use parsing as sequence labeling as a common framework to learn across constituency and dependency syntactic abstractions.To do so, we cast the problem as multitask learning (MTL).First, we show that adding a parsing paradigm as an auxiliary loss consistently improves the performance on the other paradigm.Secondly, we explore an MTL sequence labeling model that parses both representations, at almost no cost in terms of performance and speed.The results across the board show that on average MTL models with auxiliary losses for constituency parsing outperform singletask ones by 1.05 F1 points, and for dependency parsing by 0.62 UAS points.
Michalina Strzyz, David Vilares 0001, Carlos Gómez-Rodríguez
ACL (1)3
2019 HEAD-QA: A Healthcare Dataset for Complex Reasoning
abstract
We present HEAD-QA, a multi-choice question answering testbed to encourage research on complex reasoning.The questions come from exams to access a specialized position in the Spanish healthcare system, and are challenging even for highly specialized humans.We then consider monolingual (Spanish) and cross-lingual (to English) experiments with information retrieval and neural techniques.We show that: (i) HEAD-QA challenges current methods, and (ii) the results lag well behind human performance, demonstrating its usefulness as a benchmark for future work.
David Vilares 0001, Carlos Gómez-Rodríguez
ACL (1)2
2019 Speeding up Natural Language Parsing by Reusing Partial Results
Michalina Strzyz, Carlos Gómez-Rodríguez
CICLing (1)2
2019 Towards Making a Dependency Parser See
abstract
Michalina Strzyz, David Vilares, Carlos Gómez-Rodríguez. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Michalina Strzyz, David Vilares 0001, Carlos Gómez-Rodríguez
EMNLP/IJCNLP (1)3
2019 Faster shift-reduce constituent parsing with a non-binary, bottom-up strategy
Daniel Fernández-González, Carlos Gómez-Rodríguez
Artif. Intell.2
2019 Comparing neural- and N-gram-based language models for word segmentation
abstract
Word segmentation is the task of inserting or deleting word boundary characters in order to separate character sequences that correspond to words in some language. In this article we propose an approach based on a beam search algorithm and a language model working at the byte/character level, the latter component implemented either as an n-gram model or a recurrent neural network. The resulting system analyzes the text input with no word boundaries one token at a time, which can be a character or a byte, and uses the information gathered by the language model to determine if a boundary must be placed in the current position or not. Our aim is to use this system in a preprocessing step for a microtext normalization system. This means that it needs to effectively cope with the data sparsity present on this kind of texts. We also strove to surpass the performance of two readily available word segmentation systems: The well-known and accessible Word Breaker by Microsoft, and the Python module WordSegment by Grant Jenks. The results show that we have met our objectives, and we hope to continue to improve both the precision and the efficiency of our system in the future.
Yerai Doval, Carlos Gómez-Rodríguez
J. Assoc. Inf. Sci. Technol.2
2018 Global Transition-based Non-projective Dependency Parsing
abstract
Shi, Huang, and Lee (2017a) obtained state-of-the-art results for English and Chinese dependency parsing by combining dynamic-programming implementations of transition-based dependency parsers with a minimal set of bidirectional LSTM features.However, their results were limited to projective parsing.In this paper, we extend their approach to support non-projectivity by providing the first practical implementation of the MH 4 algorithm, an Opn 4 q mildly nonprojective dynamic-programming parser with very high coverage on non-projective treebanks.To make MH 4 compatible with minimal transition-based feature sets, we introduce a transition-based interpretation of it in which parser items are mapped to sequences of transitions.We thus obtain the first implementation of global decoding for non-projective transition-based parsing, and demonstrate empirically that it is more effective than its projective counterpart in parsing a number of highly non-projective languages.
Carlos Gómez-Rodríguez, Tianze Shi, Lillian Lee
ACL (1)1
2018 Dynamic Oracles for Top-Down and In-Order Shift-Reduce Constituent Parsing
abstract
We introduce novel dynamic oracles for training two of the most accurate known shiftreduce algorithms for constituent parsing: the top-down and in-order transition-based parsers.In both cases, the dynamic oracles manage to notably increase their accuracy, in comparison to that obtained by performing classic static training.In addition, by improving the performance of the state-of-the-art in-order shift-reduce parser, we achieve the best accuracy to date (92.0 F1) obtained by a fullysupervised single-model greedy shift-reduce constituent parser on the WSJ benchmark.
Daniel Fernández-González, Carlos Gómez-Rodríguez
EMNLP2
2018 Constituent Parsing as Sequence Labeling
abstract
We introduce a method to reduce constituent parsing to sequence labeling.For each word w t , it generates a label that encodes: (1) the number of ancestors in the tree that the words w t and w t+1 have in common, and (2) the nonterminal symbol at the lowest common ancestor.We first prove that the proposed encoding function is injective for any tree without unary branches.In practice, the approach is made extensible to all constituency trees by collapsing unary branches.We then use the PTB and CTB treebanks as testbeds and propose a set of fast baselines.We achieve 90.7% F-score on the PTB test set, outperforming the Vinyals et al. (2015) sequence-to-sequence parser.In addition, sacrificing some accuracy, our approach achieves the fastest constituent parsing speeds reported to date on PTB by a wide margin. 1
Carlos Gómez-Rodríguez, David Vilares 0001
EMNLP1
2018 New treebank or repurposed? On the feasibility of cross-lingual parsing of Romance languages with Universal Dependencies
abstract
Abstract This paper addresses the feasibility of cross-lingual parsing with Universal Dependencies (UD) between Romance languages, analyzing its performance when compared to the use of manually annotated resources of the target languages. Several experiments take into account factors such as the lexical distance between the source and target varieties, the impact of delexicalization, the combination of different source treebanks or the adaptation of resources to the target language, among others. The results of these evaluations show that the direct application of a parser from one Romance language to another reaches similar labeled attachment score (LAS) values to those obtained with a manual annotation of about 3,000 tokens in the target language, and unlabeled attachment score (UAS) results equivalent to the use of around 7,000 tokens, depending on the case. These numbers can noticeably increase by performing a focused selection of the source treebanks. Furthermore, the removal of the words in the training corpus (delexicalization) is not useful in most cases of cross-lingual parsing of Romance languages. The lessons learned with the performed experiments were used to build a new UD treebank for Galician, with 1,000 sentences manually corrected after an automatic cross-lingual annotation. Several evaluations in this new resource show that a cross-lingual parser built with the best combination and adaptation of the source treebanks performs better (77 percent LAS and 82 percent UAS) than using more than 16,000 (for LAS results) and more than 20,000 (UAS) manually labeled tokens of Galician.
Marcos García 0001, Carlos Gómez-Rodríguez, Miguel A. Alonso 0001
Nat. Lang. Eng.2
2017 A Full Non-Monotonic Transition System for Unrestricted Non-Projective Parsing
abstract
Restricted non-monotonicity has been shown beneficial for the projective arceager dependency parser in previous research, as posterior decisions can repair mistakes made in previous states due to the lack of information.In this paper, we propose a novel, fully non-monotonic transition system based on the non-projective Covington algorithm.As a non-monotonic system requires exploration of erroneous actions during the training process, we develop several non-monotonic variants of the recently defined dynamic oracle for the Covington parser, based on tight approximations of the loss.Experiments on datasets from the CoNLL-X and CoNLL-XI shared tasks show that a non-monotonic dynamic oracle outperforms the monotonic version in the majority of languages.
Daniel Fernández-González, Carlos Gómez-Rodríguez
ACL (1)2
2017 Generic Axiomatization of Families of Noncrossing Graphs in Dependency Parsing
abstract
We present a simple encoding for unlabeled noncrossing graphs and show how its latent counterpart helps us to represent several families of directed and undirected graphs used in syntactic and semantic parsing of natural language as contextfree languages.The families are separated purely on the basis of forbidden patterns in latent encoding, eliminating the need to differentiate the families of non-crossing graphs in inference algorithms: one algorithm works for all when the search space can be controlled in parser input.
Anssi Yli-Jyrä, Carlos Gómez-Rodríguez
ACL (1)2
2017 Supervised sentiment analysis in multilingual environments
David Vilares 0001, Miguel A. Alonso 0001, Carlos Gómez-Rodríguez
Inf. Process. Manag.3
2017 Universal, unsupervised (rule-based), uncovered sentiment analysis
David Vilares 0001, Carlos Gómez-Rodríguez, Miguel A. Alonso 0001
Knowl. Based Syst.2
2016 EN-ES-CS: An English-Spanish Code-Switching Twitter Corpus for Multilingual Sentiment Analysis
David Vilares 0001, Miguel A. Alonso 0001, Carlos Gómez-Rodríguez
LREC3
2016 Restricted Non-Projectivity: Coverage vs. Efficiency
abstract
In the last decade, various restricted classes of non-projective dependency trees have been proposed with the goal of achieving a good tradeoff between parsing efficiency and coverage of the syntactic structures found in natural languages. We perform an extensive study measuring the coverage of a wide range of such classes on corpora of 30 languages under two different syntactic annotation criteria. The results show that, among the currently known relaxations of projectivity, the best tradeoff between coverage and computational complexity of exact parsing is achieved by either 1-endpoint-crossing trees or MHktrees, depending on the level of coverage desired. We also present some properties of the relation of MHktrees to other relevant classes of trees.
Carlos Gómez-Rodríguez
Comput. Linguistics1
2015 Undirected Dependency Parsing
abstract
Dependency parsers, which are widely used in natural language processing tasks, employ a representation of syntax in which the structure of sentences is expressed in the form of directed links (dependencies) between their words. In this article, we introduce a new approach to transition‐based dependency parsing in which the parsing algorithm does not directly construct dependencies, but rather undirected links, which are then assigned a direction in a postprocessing step. We show that this alleviates error propagation, because undirected parsers do not need to observe the single‐head constraint, resulting in better accuracy. Undirected parsers can be obtained by transforming existing directed transition‐based parsers as long as they satisfy certain conditions. We apply this approach to obtain undirected variants of three different parsers (the Planar, 2‐Planar, and Covington algorithms) and perform experiments on several data sets from the CoNLL‐X shared tasks and on the Wall Street Journal portion of the Penn Treebank, showing that our approach is successful in reducing error propagation and produces improvements in parsing accuracy in most of the cases and achieving results competitive with state‐of‐the‐art transition‐based parsers.
Carlos Gómez-Rodríguez, Daniel Fernández-González, Victor M. Darriba
Comput. Intell.1
2015 On the usefulness of lexical and syntactic processing in polarity classification of Twitter messages
abstract
Millions of micro texts are published every day on Twitter. Identifying the sentiment present in them can be helpful for measuring the frame of mind of the public, their satisfaction with respect to a product, or their support of a social event. In this context, polarity classification is a subfield of sentiment analysis focused on determining whether the content of a text is objective or subjective, and in the latter case, if it conveys a positive or a negative opinion. Most polarity detection techniques tend to take into account individual terms in the text and even some degree of linguistic knowledge, but they do not usually consider syntactic relations between words. This article explores how relating lexical, syntactic, and psychometric information can be helpful to perform polarity classification on Spanish tweets. We provide an evaluation for both shallow and deep linguistic perspectives. Empirical results show an improved performance of syntactic approaches over pure lexical models when using large training sets to create a classifier, but this tendency is reversed when small training collections are used.
David Vilares 0001, Miguel A. Alonso 0001, Carlos Gómez-Rodríguez
J. Assoc. Inf. Sci. Technol.3
2015 A syntactic approach for opinion mining on Spanish reviews
abstract
Abstract We describe an opinion mining system which classifies the polarity of Spanish texts. We propose an NLP approach that undertakes pre-processing, tokenisation and POS tagging of texts to then obtain the syntactic structure of sentences by means of a dependency parser. This structure is then used to address three of the most significant linguistic constructions for the purpose in question: intensification, subordinate adversative clauses and negation. We also propose a semi-automatic domain adaptation method to improve the accuracy of our system in specific application domains, by enriching semantic dictionaries using machine learning methods in order to adapt the semantic orientation of their words to a particular field. Experimental results are promising in both general and specific domains.
David Vilares 0001, Miguel A. Alonso 0001, Carlos Gómez-Rodríguez
Nat. Lang. Eng.3
2014 A Polynomial-Time Dynamic Oracle for Non-Projective Dependency Parsing
abstract
The introduction of dynamic oracles has considerably improved the accuracy of greedy transition-based dependency parsers, without sacrificing parsing efficiency.However, this enhancement is limited to projective parsing, and dynamic oracles have not yet been implemented for parsers supporting non-projectivity.In this paper we introduce the first such oracle, for a non-projective parser based on Attardi's parser.We show that training with this oracle improves parsing accuracy over a conventional (static) oracle on a wide range of datasets.
Carlos Gómez-Rodríguez, Francesco Sartorio, Giorgio Satta
EMNLP1
2014 Finding the smallest binarization of a CFG is NP-hard
Carlos Gómez-Rodríguez
J. Comput. Syst. Sci.1
2013 Supervised polarity classification of Spanish tweets based on linguistic knowledge
abstract
We describe a system that classifies the polarity of Spanish tweets. We adopt a hybrid approach, which combines machine learning and linguistic knowledge acquired by means of NLP. We use part-of-speech tags, syntactic dependencies and semantic knowledge as features for a supervised classifier. Lexical particularities of the language used in Twitter are taken into account in a pre-processing step. Experimental results improve over those of pure machine learning approaches and confirm the practical utility of the proposal.
David Vilares 0001, Miguel A. Alonso 0001, Carlos Gómez-Rodríguez
ACM Symposium on Document Engineering3
2013 Divisible Transition Systems and Multiplanar Dependency Parsing
abstract
Transition-based parsing is a widely used approach for dependency parsing that combines high efficiency with expressive feature models. Many different transition systems have been proposed, often formalized in slightly different frameworks. In this article, we show that a large number of the known systems for projective dependency parsing can be viewed as variants of the same stack-based system with a small set of elementary transitions that can be composed into complex transitions and restricted in different ways. We call these systems divisible transition systems and prove a number of theoretical results about their expressivity and complexity. In particular, we characterize an important subclass called efficient divisible transition systems that parse planar dependency graphs in linear time. We go on to show, first, how this system can be restricted to capture exactly the set of planar dependency trees and, secondly, how the system can be generalized to k-planar trees by making use of multiple stacks. Using the first known efficient test for k-planarity, we investigate the coverage of k-planar trees in available dependency treebanks and find a very good fit for 2-planar trees. We end with an experimental evaluation showing that our 2-planar parser gives significant improvements in parsing accuracy over the corresponding 1-planar and projective parsers for data sets with non-projective dependency trees and performs on a par with the widely used arc-eager pseudo-projective parser.
Carlos Gómez-Rodríguez, Joakim Nivre
Comput. Linguistics1
2012 Improving Transition-Based Dependency Parsing with Buffer Transitions
Daniel Fernández-González, Carlos Gómez-Rodríguez
EMNLP-CoNLL2
2011 Dynamic Programming Algorithms for Transition-Based Dependency Parsers
Marco Kuhlmann, Carlos Gómez-Rodríguez, Giorgio Satta
ACL2
2011 Exact Inference for Generative Probabilistic Non-Projective Dependency Parsing
Shay B. Cohen, Carlos Gómez-Rodríguez, Giorgio Satta
EMNLP2
2011 Dependency Parsing Schemata and Mildly Non-Projective Dependency Parsing
abstract
We introduce dependency parsing schemata, a formal framework based on Sikkel's parsing schemata for constituency parsers, which can be used to describe, analyze, and compare dependency parsing algorithms. We use this framework to describe several well-known projective and non-projective dependency parsers, build correctness proofs, and establish formal relationships between them. We then use the framework to define new polynomial-time parsing algorithms for various mildly non-projective dependency formalisms, including well-nested structures with their gap degree bounded by a constant k in time O(n5+2k), and a new class that includes all gap degree k structures present in several natural language treebanks (which we call mildly ill-nested structures for gap degree k) in time O(n4+3k). Finally, we illustrate how the parsing schema framework can be applied to Link Grammar, a dependency-related formalism.
Carlos Gómez-Rodríguez, John Carroll 0001, David J. Weir
Comput. Linguistics1
2010 A Transition-Based Parser for 2-Planar Dependency Structures
Carlos Gómez-Rodríguez, Joakim Nivre
ACL1
2010 Evaluation of Dependency Parsers on Unbounded Dependencies
Joakim Nivre, Laura Rimell, Ryan T. McDonald, Carlos Gómez-Rodríguez
COLING4
2010 Efficient Parsing of Well-Nested Linear Context-Free Rewriting Systems
Carlos Gómez-Rodríguez, Marco Kuhlmann, Giorgio Satta
HLT-NAACL1
2010 Error-repair parsing schemata
Carlos Gómez-Rodríguez, Miguel A. Alonso 0001, Manuel Vilares Ferro
Theor. Comput. Sci.1
2009 An Optimal-Time Binarization Algorithm for Linear Context-Free Rewriting Systems with Fan-Out Two
Carlos Gómez-Rodríguez, Giorgio Satta
ACL/IJCNLP1
2009 A General Method for Transforming Standard Parsers into Error-Repair Parsers
Carlos Gómez-Rodríguez, Miguel A. Alonso 0001, Manuel Vilares Ferro
CICLing1
2009 Parsing Mildly Non-Projective Dependency Structures
Carlos Gómez-Rodríguez, David J. Weir, John Carroll 0001
EACL1
2009 Optimal Reduction of Rule Length in Linear Context-Free Rewriting Systems
Carlos Gómez-Rodríguez, Marco Kuhlmann, Giorgio Satta, David J. Weir
HLT-NAACL1
2009 A compiler for parsing schemata
abstract
Abstract We present a compiler that can be used to automatically obtain efficient Java implementations of parsing algorithms from formal specifications expressed as parsing schemata. The system performs an analysis of the inference rules in the input schemata in order to determine the best data structures and indexes to use, and to ensure that the generated implementations are efficient. The system described is general enough to be able to handle all kinds of schemata for different grammar formalisms, such as context‐free grammars and tree‐adjoining grammars, and it provides an extensibility mechanism allowing the user to define custom notational elements. This compiler has proven very useful for analyzing, prototyping and comparing natural‐language parsers in real domains, as can be seen in the empirical examples provided at the end of the paper. Copyright © 2008 John Wiley & Sons, Ltd.
Carlos Gómez-Rodríguez, Jesús Vilares, Miguel A. Alonso 0001
Softw. Pract. Exp.1
2008 A Deductive Approach to Dependency Parsing
Carlos Gómez-Rodríguez, John Carroll 0001, David J. Weir
ACL1
2007 Compiling Declarative Specifications of Parsing Algorithms
Carlos Gómez-Rodríguez, Jesús Vilares, Miguel A. Alonso 0001
DEXA1
2005 Managing syntactic variation in text retrieval
abstract
Information Retrieval systems are limited by the linguistic variation of language. The use of Natural Language Processing techniques to manage this problem has been studied for a long time, but mainly focusing on English. In this paper we deal with European languages, taking Spanish as a case in point. Two different sources of syntactic information, queries and documents, are studied in order to increase the performance of Information Retrieval systems.
Jesús Vilares, Carlos Gómez-Rodríguez, Miguel A. Alonso 0001
ACM Symposium on Document Engineering2