Xavier Carreras

dblp:46/4131 · DBLP profile ↗
← Back
45ranked-venue papers
17as first author
0since 2021 · last 2019
0000-0001-7432-4540ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 43 · 16 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Information extraction and text analysis · 35% Language models and text generation · 21% Probabilistic and Bayesian machine learning · 20%
Theoretical computer science
6 papers
Automata and formal languages · 88% Mathematical optimization · 7% Algorithms and data structures · 4%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Automata and formal languages › weighted automata
spectral learning
0.942019
Interpolated Spectral NGram Language Models · ACL (1) 2019
Unsupervised Spectral Learning of Finite State Transducers · NIPS 2013
Unsupervised Spectral Learning of WCFG as Low-rank Matrix Completion · EMNLP 2013
Automata and formal languages
weighted automata
0.522019
Interpolated Spectral NGram Language Models · ACL (1) 2019
Local Loss Optimization in Operator Models: A New Insight into Spectral Learning · ICML 2012
Natural language and speech › Language models and text generation › language modeling
n-gram language model
0.412019
Interpolated Spectral NGram Language Models · ACL (1) 2019
Natural language and speech › Language models and text generation › language modeling
statistical language modeling
0.412019
Interpolated Spectral NGram Language Models · ACL (1) 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.322013
Unsupervised Spectral Learning of Finite State Transducers · NIPS 2013
Local Loss Optimization in Operator Models: A New Insight into Spectral Learning · ICML 2012
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.342014
Non-Projective Parsing for Statistical Machine Translation · EMNLP 2009
Experiments with a Higher-Order Projective Dependency Parser · EMNLP-CoNLL 2007
A Shortest-path Method for Arc-factored Semantic Role Labeling · EMNLP 2014
Natural language and speech › Information extraction and text analysis › named entity recognition
named entity classification
0.212015
Low-Rank Regularization for Sparse Conjunctive Feature Spaces: An Application to Named Entity Classification · ACL (1) 2015
Natural language and speech › Information extraction and text analysis
named entity recognition
0.212015
Named entity recognition with document-specific KB tag gazetteers · EMNLP 2015
Natural language and speech › Information extraction and text analysis
semantic role labeling
0.212014
A Shortest-path Method for Arc-factored Semantic Role Labeling · EMNLP 2014
Machine learning › Deep learning architectures and training › regularization
spectral regularization
0.212014
Spectral Regularization for Max-Margin Sequence Tagging · ICML 2014
Machine learning and data management
structured prediction
0.212014
Spectral Regularization for Max-Margin Sequence Tagging · ICML 2014
Automata and formal languages › formal grammars
context-free grammar
0.212013
Unsupervised Spectral Learning of WCFG as Low-rank Matrix Completion · EMNLP 2013
Automata and formal languages › transducers
finite-state transducers
0.212013
Unsupervised Spectral Learning of Finite State Transducers · NIPS 2013
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
0.222008
Simple Semi-supervised Dependency Parsing · ACL 2008
Experiments with a Higher-Order Projective Dependency Parser · EMNLP-CoNLL 2007
Machine learning › Optimization for machine learning › mirror descent
exponentiated gradient
0.222008
Exponentiated Gradient Algorithms for Conditional Random Fields and Max-Margin Markov Networks · J. Mach. Learn. Res. 2008
Exponentiated gradient algorithms for log-linear structured prediction · ICML 2007
Machine learning › Probabilistic and Bayesian machine learning
structured prediction
0.222008
Exponentiated Gradient Algorithms for Conditional Random Fields and Max-Margin Markov Networks · J. Mach. Learn. Res. 2008
Exponentiated gradient algorithms for log-linear structured prediction · ICML 2007
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
non-projective dependency parsing
0.112009
Non-Projective Parsing for Statistical Machine Translation · EMNLP 2009
Machine learning › Deep learning architectures and training
regularization
0.112009
An efficient projection for l1,infinity regularization · ICML 2009
Natural language and speech › Machine translation
statistical machine translation
0.112009
Non-Projective Parsing for Statistical Machine Translation · EMNLP 2009
Mathematical optimization › gradient descent
projected gradient descent
0.112009
An efficient projection for l1,infinity regularization · ICML 2009
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
conditional random field
0.112008
Exponentiated Gradient Algorithms for Conditional Random Fields and Max-Margin Markov Networks · J. Mach. Learn. Res. 2008
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
max-margin markov networks
0.112008
Exponentiated Gradient Algorithms for Conditional Random Fields and Max-Margin Markov Networks · J. Mach. Learn. Res. 2008
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
semi-supervised dependency parsing
0.112008
Simple Semi-supervised Dependency Parsing · ACL 2008
Machine learning › Optimization for machine learning
convex optimization
0.112007
Exponentiated gradient algorithms for log-linear structured prediction · ICML 2007
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › exponential family
maximum entropy models
0.112007
Exponentiated gradient algorithms for log-linear structured prediction · ICML 2007
Algorithms and data structures › learning algorithms
structured prediction
0.112007
Structured Prediction Models via the Matrix-Tree Theorem · EMNLP-CoNLL 2007
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge base
0.112015
Named entity recognition with document-specific KB tag gazetteers · EMNLP 2015
Machine learning › Deep learning architectures and training › regularization
low-rank regularization
0.112015
Low-Rank Regularization for Sparse Conjunctive Feature Spaces: An Application to Named Entity Classification · ACL (1) 2015
Mathematical optimization › continuous optimization › convex optimization › norm optimization
nuclear norm minimization
0.012013
Unsupervised Spectral Learning of Finite State Transducers · NIPS 2013
Machine learning › Learning theory
online learning
0.012003
Online Learning via Global Feedback for Phrase Recognition · NIPS 2003

Methods — techniques the papers use, named apart from their topics

spectral learning · 1.4moment matching · 0.8interpolation · 0.8convex optimization · 0.5observable operator models · 0.4sparse conjunctive features · 0.2low-rank regularization · 0.2knowledge base tags · 0.2gazetteer · 0.2spectral regularization · 0.2shortest-path inference · 0.2rank minimization · 0.2nuclear norm · 0.2low-rank matrix completion · 0.2convex relaxation · 0.2local loss optimization · 0.1projected gradient · 0.1joint sparsity · 0.1
YearPublicationVenuePosition
2019 Interpolated Spectral NGram Language Models
abstract
Spectral models for learning weighted nondeterministic automata have nice theoretical and algorithmic properties.Despite this, it has been challenging to obtain competitive results in language modeling tasks, for two main reasons.First, in order to capture long-range dependencies of the data, the method must use statistics from long substrings, which results in very large matrices that are difficult to decompose.The second is that the loss function behind spectral learning, based on moment matching, differs from the probabilistic metrics used to evaluate language models.In this work we employ a technique for scaling up spectral learning, and use interpolated predictions that are optimized to maximize perplexity.Our experiments in character-based language modeling show that our method matches the performance of stateof-the-art ngram models, while being very fast to train.
Ariadna Quattoni, Xavier Carreras
ACL (1)2
2018 Local String Transduction as Sequence Labeling
abstract
We show that the general problem of string transduction can be reduced to the problem of sequence labeling. While character deletion and insertions are allowed in string transduction, they do not exist in sequence labeling. We show how to overcome this difference. Our approach can be used with any sequence labeling algorithm and it works best for problems in which string transduction imposes a strong notion of locality (no long range dependencies). We experiment with spelling correction for social media, OCR correction, and morphological inflection, and we see that it behaves better than seq2seq models and yields state-of-the-art results in several cases.
Joana Ribeiro, Shashi Narayan, Shay B. Cohen, Xavier Carreras
COLING4
2017 A Maximum Matching Algorithm for Basis Selection in Spectral Learning
abstract
We present a solution to scale spectral algorithms for learning sequence functions. We are interested in the case where these functions are sparse (that is, for most sequences they return 0). Spectral algorithms reduce the learning problem to the task of computing an SVD decomposition over a special type of matrix called the Hankel matrix. This matrix is designed to capture the relevant statistics of the training sequences. What is crucial is that to capture long range dependencies we must consider very large Hankel matrices. Thus the computation of the SVD becomes a critical bottleneck. Our solution finds a subset of rows and columns of the Hankel that realizes a compact and informative Hankel submatrix. The novelty lies in the way that this subset is selected: we exploit a maximal bipartite matching combinatorial algorithm to look for a sub-block with full structural rank, and show how computation of this sub-block can be further improved by exploiting the specific structure of Hankel matrices.
Ariadna Quattoni, Xavier Carreras, Matthias Gallé
AISTATS2
2015 Low-Rank Regularization for Sparse Conjunctive Feature Spaces: An Application to Named Entity Classification
abstract
Audi Primadhanty, Xavier Carreras, Ariadna Quattoni. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Audi Primadhanty, Xavier Carreras, Ariadna Quattoni
ACL (1)2
2015 Transition-based Spinal Parsing
abstract
We present a transition-based arc-eager model to parse spinal trees, a dependencybased representation that includes phrasestructure information in the form of constituent spines assigned to tokens.As a main advantage, the arc-eager model can use a rich set of features combining dependency and constituent information, while parsing in linear time.We describe a set of conditions for the arc-eager system to produce valid spinal structures.In experiments using beam search we show that the model obtains a good trade-off between speed and accuracy, and yields state of the art performance for both dependency and constituent parsing measures.
Miguel Ballesteros, Xavier Carreras
CoNLL2
2015 Named entity recognition with document-specific KB tag gazetteers
abstract
We consider a novel setting for Named Entity Recognition (NER) where we have access to document-specific knowledge base tags.These tags consist of a canonical name from a knowledge base (KB) and entity type, but are not aligned to the text.We explore how to use KB tags to create document-specific gazetteers at inference time to improve NER.We find that this kind of supervision helps recognise organisations more than standard widecoverage gazetteers.Moreover, augmenting document-specific gazetteers with KB information lets users specify fewer tags for the same performance, reducing cost.
Will Radford, Xavier Carreras, James Henderson 0001
EMNLP2
2014 Learning Task-specific Bilexical Embeddings
Pranava Swaroop Madhyastha, Xavier Carreras, Ariadna Quattoni
COLING2
2014 XLike Project Language Analysis Services
abstract
Xavier Carreras, Lluís Padró, Lei Zhang, Achim Rettinger, Zhixing Li, Esteban García-Cuesta, Željko Agić, Božo Bekavac, Blaz Fortuna, Tadej Štajner. Proceedings of the Demonstrations at the 14th Conference of the European Chapter of the Association for Computational Linguistics. 2014.
Xavier Carreras, Lluís Padró 0001, Lei Zhang 0034, Achim Rettinger, Esteban García-Cuesta, Zeljko Agic, Bozo Bekavac, Blaz Fortuna, Tadej Stajner
EACL1
2014 A Shortest-path Method for Arc-factored Semantic Role Labeling
abstract
We introduce a Semantic Role Labeling (SRL) parser that finds semantic roles for a predicate together with the syntactic paths linking predicates and arguments.Our main contribution is to formulate SRL in terms of shortest-path inference, on the assumption that the SRL model is restricted to arc-factored features of the syntactic paths behind semantic roles.Overall, our method for SRL is a novel way to exploit larger variability in the syntactic realizations of predicate-argument relations, moving away from pipeline architectures.Experiments show that our approach improves the robustness of the predictions, producing arc-factored models that perform closely to methods using unrestricted features from the syntax.
Xavier Lluís, Xavier Carreras, Lluís Màrquez
EMNLP2
2014 Spectral Regularization for Max-Margin Sequence Tagging
abstract
We frame max-margin learning of latent variable structured prediction models as a convex optimization problem, making use of scoring functions computed by input-output observable operator models. This learning problem can be expressed as an optimization involving a low-rank Hankel matrix that represents the input-output operator model. The direct outcome of our work is a new spectral regularization method for max-margin structured prediction. Our experiments confirm that our proposed regularization framework leads to an effective way of controlling the capacity of structured prediction models.
Ariadna Quattoni, Borja Balle, Xavier Carreras, Amir Globerson
ICML3
2014 Language Processing Infrastructure in the XLike Project
Lluís Padró 0001, Zeljko Agic, Xavier Carreras, Blaz Fortuna, Esteban García-Cuesta, Tadej Stajner, Marko Tadic
LREC3
2014 Spectral learning of weighted automata - A forward-backward perspective
Borja Balle, Xavier Carreras, Franco M. Luque, Ariadna Quattoni
Mach. Learn.2
2013 Unsupervised Spectral Learning of WCFG as Low-rank Matrix Completion
abstract
We derive a spectral method for unsupervised learning of Weighted Context Free Grammars.We frame WCFG induction as finding a Hankel matrix that has low rank and is linearly constrained to represent a function computed by inside-outside recursions.The proposed algorithm picks the grammar that agrees with a sample and is the simplest with respect to the nuclear norm of the Hankel matrix.
Raphaël Bailly, Xavier Carreras, Franco M. Luque, Ariadna Quattoni
EMNLP2
2013 Unsupervised Spectral Learning of Finite State Transducers
abstract
Finite-State Transducers (FST) are a standard tool for modeling paired input-output sequences and are used in numerous applications, ranging from computational biology to natural language processing. Recently Balle et al. presented a spectral algorithm for learning FST from samples of aligned input-output sequences. In this paper we address the more realistic, yet challenging setting where the alignments are unknown to the learning algorithm. We frame FST learning as finding a low rank Hankel matrix satisfying constraints derived from observable statistics. Under this formulation, we provide identifiability results for FST distributions. Then, following previous work on rank minimization, we propose a regularized convex relaxation of this objective which is based on minimizing a nuclear norm penalty subject to linear constraints and can be solved efficiently.
Raphaël Bailly, Xavier Carreras, Ariadna Quattoni
NIPS2
2013 Joint Arc-factored Parsing of Syntactic and Semantic Dependencies
abstract
In this paper we introduce a joint arc-factored model for syntactic and semantic dependency parsing. The semantic role labeler predicts the full syntactic paths that connect predicates with their arguments. This process is framed as a linear assignment task, which allows to control some well-formedness constraints. For the syntactic part, we define a standard arc-factored dependency model that predicts the full syntactic tree. Finally, we employ dual decomposition techniques to produce consistent syntactic and predicate-argument structures while searching over a large space of syntactic configurations. In experiments on the CoNLL-2009 English benchmark we observe very competitive results.
Xavier Lluís, Xavier Carreras, Lluís Màrquez
Trans. Assoc. Comput. Linguistics2
2012 Spectral Learning for Non-Deterministic Dependency Parsing
Franco M. Luque, Ariadna Quattoni, Borja Balle, Xavier Carreras
EACL4
2012 A Latent Variable Ranking Model for Content-Based Retrieval
Ariadna Quattoni, Xavier Carreras, Antonio Torralba 0001
ECIR2
2012 Local Loss Optimization in Operator Models: A New Insight into Spectral Learning
Borja Balle, Ariadna Quattoni, Xavier Carreras
ICML3
2011 A Spectral Learning Algorithm for Finite State Transducers
Borja Balle, Ariadna Quattoni, Xavier Carreras
ECML/PKDD (1)3
2009 Non-Projective Parsing for Statistical Machine Translation
Xavier Carreras, Michael Collins 0001
EMNLP1
2009 An Empirical Study of Semi-supervised Structured Conditional Models for Dependency Parsing
Jun Suzuki 0001, Hideki Isozaki, Xavier Carreras, Michael Collins 0001
EMNLP3
2009 An efficient projection for l1,infinity regularization
abstract
In recent years the l1, ∞ norm has been proposed for joint regularization. In essence, this type of regularization aims at extending the l1 framework for learning sparse models to a setting where the goal is to learn a set of jointly sparse models. In this paper we derive a simple and effective projected gradient method for optimization of l1, ∞ regularized problems. The main challenge in developing such a method resides on being able to compute efficient projections to the l1, ∞ ball. We present an algorithm that works in O(n log n) time and O(n) memory where n is the number of parameters. We test our algorithm in a multi-task image annotation problem. Our results show that l1, ∞ leads to better performance than both l2 and l1 regularization and that it is is effective in discovering jointly sparse solutions.
Ariadna Quattoni, Xavier Carreras, Michael Collins 0001, Trevor Darrell
ICML2
2008 Simple Semi-supervised Dependency Parsing
Terry Koo, Xavier Carreras, Michael Collins 0001
ACL2
2008 TAG, Dynamic Programming, and the Perceptron for Efficient, Feature-Rich Parsing
Xavier Carreras, Michael Collins 0001, Terry Koo
CoNLL1
2008 Semantic Role Labeling: An Introduction to the Special Issue
abstract
Semantic role labeling, the computational identification and labeling of arguments in text, has become a leading task in computational linguistics today. Although the issues for this task have been studied for decades, the availability of large resources and the development of statistical machine learning methods have heightened the amount of effort in this field. This special issue presents selected and representative work in the field. This overview describes linguistic background of the problem, the movement from linguistic theories to computational practice, the major resources that are being used, an overview of steps taken in computational systems, and a description of the key issues and results in semantic role labeling (as revealed in several international evaluations). We assess weaknesses in semantic role labeling and identify important challenges facing the field. Overall, the opportunities and the potential for useful further research in semantic role labeling are considerable.
Lluís Màrquez, Xavier Carreras, Kenneth C. Litkowski, Suzanne Stevenson
Comput. Linguistics2
2008 Exponentiated Gradient Algorithms for Conditional Random Fields and Max-Margin Markov Networks
Michael Collins 0001, Amir Globerson, Terry Koo, Xavier Carreras, Peter L. Bartlett
J. Mach. Learn. Res.4
2007 Experiments with a Higher-Order Projective Dependency Parser
Xavier Carreras
EMNLP-CoNLL1
2007 Structured Prediction Models via the Matrix-Tree Theorem
Terry Koo, Amir Globerson, Xavier Carreras, Michael Collins 0001
EMNLP-CoNLL3
2007 Exponentiated gradient algorithms for log-linear structured prediction
abstract
Conditional log-linear models are a commonly used method for structured prediction. Efficient learning of parameters in these models is therefore an important problem. This paper describes an exponentiated gradient (EG) algorithm for training such models. EG is applied to the convex dual of the maximum likelihood objective; this results in both sequential and parallel update algorithms, where in the sequential algorithm parameters are updated in an online fashion. We provide a convergence proof for both algorithms. Our analysis also simplifies previous results on EG for max-margin models, and leads to a tighter bound on convergence rates. Experiments on a large-scale parsing task show that the proposed algorithm converges much faster than conjugate-gradient and L-BFGS approaches both in terms of optimization objective and test error.
Amir Globerson, Terry Koo, Xavier Carreras, Michael Collins 0001
ICML3
2007 Combination Strategies for Semantic Role Labeling
abstract
This paper introduces and analyzes a battery of inference models for the problem of semantic role labeling: one based on constraint satisfaction, and several strategies that model the inference as a meta-learning problem using discriminative classifiers. These classifiers are developed with a rich set of novel features that encode proposition and sentence-level information. To our knowledge, this is the first work that: (a) performs a thorough analysis of learning-based inference models for semantic role labeling, and (b) compares several inference strategies in this context. We evaluate the proposed inference strategies in the framework of the CoNLL-2005 shared task using only automatically-generated syntactic information. The extensive experimental evaluation and analysis indicates that all the proposed inference strategies are successful -they all outperform the current best results reported in the CoNLL-2005 evaluation exercise- but each of the proposed approaches has its advantages and disadvantages. Several important traits of a state-of-the-art SRL combination strategy emerge from this analysis: (i) individual models should be combined at the granularity of candidate arguments rather than at the granularity of complete solutions; (ii) the best combination strategy uses an inference model based in learning; and (iii) the learning-based inference benefits from max-margin classifiers and global feedback.
Mihai Surdeanu, Lluís Màrquez, Xavier Carreras, Pere Comas
J. Artif. Intell. Res.3
2006 Projective Dependency Parsing with Perceptron
Xavier Carreras, Mihai Surdeanu, Lluís Màrquez
CoNLL1
2005 Introduction to the CoNLL-2005 Shared Task: Semantic Role Labeling
Xavier Carreras, Lluís Màrquez
CoNLL1
2005 Filtering-Ranking Perceptron Learning for Partial Parsing
Xavier Carreras, Lluís Màrquez, C. Jorge Castro
Mach. Learn.1
2004 Introduction to the CoNLL-2004 Shared Task: Semantic Role Labeling
Xavier Carreras, Lluís Màrquez
CoNLL1
2004 Hierarchical Recognition of Propositional Arguments with Perceptrons
Xavier Carreras, Lluís Màrquez, Grzegorz Chrupala
CoNLL1
2004 Exploiting diversity of margin-based classifiers
abstract
An experimental comparison among support vector machines, Ada boost and a recently proposed model for maximizing the margin with feed-forward neural networks has been made on a real-world classification problem, namely text categorization. The results obtained when comparing their agreement on the predictions show that similar performance does not imply similar predictions, suggesting that different models can be combined to obtain better performance. As a consequence of the study, we derived a very simple confidence measure of the prediction of the tested margin-based classifiers. This measure is based on the margin curve. The combination of margin based classifiers with this confidence measure lead to a marked improvement on the performance of the system, when combined with several well-known combination schemes.
Enrique Romero, Xavier Carreras, Lluís Màrquez
IJCNN2
2004 FreeLing: An Open-Source Suite of Language Analyzers
Xavier Carreras, Isaac Chao, Lluís Padró 0001, Muntsa Padró
LREC1
2004 Margin maximization with feed-forward neural networks: a comparative study with SVM and AdaBoost
Enrique Romero, Lluís Màrquez, Xavier Carreras
Neurocomputing3
2003 A Simple Named Entity Extractor using AdaBoost
Xavier Carreras, Lluís Màrquez, Lluís Padró 0001
CoNLL1
2003 Learning a Perceptron-Based Named Entity Chunker via Online Recognition Feedback
Xavier Carreras, Lluís Màrquez, Lluís Padró 0001
CoNLL1
2003 Named Entity Recognition For Catalan Using Only Spanish Resources and Unlabelled Data
Xavier Carreras, Lluís Màrquez, Lluís Padró 0001
EACL1
2003 Online Learning via Global Feedback for Phrase Recognition
abstract
This work presents an architecture based on perceptrons to recognize phrase structures, and an online learning algorithm to train the percep- trons together and dependently. The recognition strategy applies learning in two layers: a filtering layer, which reduces the search space by identi- fying plausible phrase candidates, and a ranking layer, which recursively builds the optimal phrase structure. We provide a recognition-based feed- back rule which reflects to each local function its committed errors from a global point of view, and allows to train them together online as percep- trons. Experimentation on a syntactic parsing problem, the recognition of clause hierarchies, improves state-of-the-art results and evinces the advantages of our global training method over optimizing each function locally and independently.
Xavier Carreras, Lluís Màrquez
NIPS1
2002 Named Entity Extraction using AdaBoost
Xavier Carreras, Lluís Màrquez, Lluís Padró 0001
CoNLL1
2002 Learning and Inference for Clause Identification
Xavier Carreras, Lluís Màrquez, Vasin Punyakanok, Dan Roth 0001
ECML1
2002 A Flexible Distributed Architecture for Natural Language Analyzers
Xavier Carreras, Lluís Padró 0001
LREC1