VLDB 2026 Research / reviewers in the wild / expert
Joseph Le Roux
dblp:25/5993
· DBLP profile ↗
17ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0002-3889-8536ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Artificial intelligence
6 papers |
Information extraction and text analysis · 77% Probabilistic and Bayesian machine learning · 23% | |
| Theoretical computer science
4 papers |
Mathematical optimization · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Mathematical optimization
lagrangian relaxation |
1.0 | 2 | 2024 | Predicting Lagrangian Multipliers for Mixed Integer Linear Programs · ICML 2024 Dependency Parsing with Bounded Block Degree and Well-nestedness via Lagrangian Relaxation and Branch-and-Bound · ACL (1) 2016 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
1.0 | 1 | 2026 | A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models · SIGIR 2026 |
Information retrieval › indexing
index compression |
1.0 | 1 | 2026 | A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models · SIGIR 2026 |
Information retrieval › retrieval models › neural retrieval
late interaction retrieval |
1.0 | 1 | 2026 | A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models · SIGIR 2026 |
Information retrieval
token pruning |
1.0 | 1 | 2026 | A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models · SIGIR 2026 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.9 | 4 | 2017 | Efficient Discontinuous Phrase-Structure Parsing via the Generalized Maximum Spanning Arborescence · EMNLP 2017 Dependency Parsing with Bounded Block Degree and Well-nestedness via Lagrangian Relaxation and Branch-and-Bound · ACL (1) 2016 Foreebank: Syntactic Analysis of Customer Support Forums · EMNLP 2015 |
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
conditional random field |
0.9 | 1 | 2025 | Bregman Conditional Random Fields: Sequence Labeling with Parallelizable Inference Algorithms · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis
sequence labeling |
0.9 | 1 | 2025 | Bregman Conditional Random Fields: Sequence Labeling with Parallelizable Inference Algorithms · ACL (1) 2025 |
Mathematical optimization › continuous optimization › convex optimization
bregman projections |
0.9 | 1 | 2025 | Bregman Conditional Random Fields: Sequence Labeling with Parallelizable Inference Algorithms · ACL (1) 2025 |
Mathematical optimization › relaxation
continuous relaxation |
0.8 | 1 | 2024 | Predicting Lagrangian Multipliers for Mixed Integer Linear Programs · ICML 2024 |
Mathematical optimization › discrete optimization
mixed integer linear programming |
0.8 | 1 | 2024 | Predicting Lagrangian Multipliers for Mixed Integer Linear Programs · ICML 2024 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing |
0.7 | 3 | 2017 | Efficient Discontinuous Phrase-Structure Parsing via the Generalized Maximum Spanning Arborescence · EMNLP 2017 Dependency Parsing with Bounded Block Degree and Well-nestedness via Lagrangian Relaxation and Branch-and-Bound · ACL (1) 2016 Semi-supervised Dependency Parsing using Lexical Affinities · ACL (1) 2012 |
Information retrieval
embedding space geometry |
0.3 | 1 | 2026 | A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models · SIGIR 2026 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
non-projective dependency parsing |
0.3 | 1 | 2017 | Efficient Discontinuous Phrase-Structure Parsing via the Generalized Maximum Spanning Arborescence · EMNLP 2017 |
Empirical software engineering › mining software repositories › developer communication analysis
developer forum analysis |
0.2 | 1 | 2015 | Foreebank: Syntactic Analysis of Customer Support Forums · EMNLP 2015 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
semi-supervised dependency parsing |
0.1 | 1 | 2012 | Semi-supervised Dependency Parsing using Lexical Affinities · ACL (1) 2012 |
Mathematical optimization › large-scale optimization › decomposition methods
dual decomposition |
0.0 | 1 | 2013 | Combining PCFG-LA Models with Dual Decomposition: A Case Study with Function Labels and Binarization · EMNLP 2013 |
Methods — techniques the papers use, named apart from their topics
parallelizable inference · 1.7fenchel-young losses · 1.7bregman projections · 1.7voronoi cell estimation · 1.0hyperspace geometry · 1.0lagrangian relaxation · 0.8graph neural network · 0.8deep learning · 0.8amortized optimization · 0.8branch-and-bound · 0.5treebank annotation · 0.4error impact analysis · 0.4dual decomposition · 0.3PCFG-LA parsing · 0.3generalized maximum spanning arborescence · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval ModelsabstractLate-interaction models such as ColBERT offer competitive performance across various retrieval tasks but require storing a dense embedding for each document token, leading to a substantial index storage overhead. Past works address this by attempting to prune low-importance token embeddings based on statistical and empirical measures, but they often either lack formal grounding or are ineffective. To address these shortcomings, we introduce a framework grounded in hyperspace geometry and cast token pruning as a Voronoi cell estimation problem in the embedding space. By interpreting each token's influence as a measure of its Voronoi region, our approach enables principled pruning that retains retrieval quality while reducing index size. Through our experiments, we demonstrate that this approach serves not only as a competitive pruning strategy but also as a valuable tool for improving and interpreting token-level behavior within dense retrieval systems. Yash Kankanampati, Yuxuan Zong, Nadi Tomeh, Benjamin Piwowarski, Joseph Le Roux |
SIGIR | 5 |
| 2025 | Bregman Conditional Random Fields: Sequence Labeling with Parallelizable Inference AlgorithmsabstractWe propose a novel discriminative model for sequence labeling called Bregman conditional random fields (BCRF).Contrary to standard linear-chain conditional random fields, BCRF allows fast parallelizable inference algorithms based on iterative Bregman projections.We show how such models can be learned using Fenchel-Young losses, including extension for learning from partial labels.Experimentally, our approach delivers comparable results to CRF while being faster, and achieves better results in highly constrained settings compared to mean field, another parallelizable alternative. Caio F. Corro, Mathieu Lacroix 0001, Joseph Le Roux |
ACL (1) | 3 |
| 2024 | Predicting Lagrangian Multipliers for Mixed Integer Linear ProgramsabstractLagrangian Relaxation stands among the most efficient approaches for solving Mixed Integer Linear Programs (MILPs) with difficult constraints. Given any duals for these constraints, called Lagrangian Multipliers (LMs), it returns a bound on the optimal value of the MILP, and Lagrangian methods seek the LMs giving the best such bound. But these methods generally rely on iterative algorithms resembling gradient descent to maximize the concave piecewise linear dual function: the computational burden grows quickly with the number of relaxed constraints. We introduce a deep learning approach that bypasses the descent, effectively amortizing per instance optimization. A probabilistic encoder based on a graph neural network computes, given a MILP instance and its Continuous Relaxation (CR) solution, high-dimensional representations of relaxed constraints, which are turned into LMs by a decoder. We train the encoder and the decoder jointly by directly optimizing the bound obtained from the predicted multipliers. Our method is applicable to any problem with a compact MILP formulation, and to any Lagrangian Relaxation providing a tighter bound than CR. Experiments on two widely known problems, Multi-Commodity Network Design and Generalized Assignment, show that our approach closes up to 85% of the gap between the continuous relaxation and the best Lagrangian bound, and provides a high-quality warm-start for descent-based Lagrangian methods. Francesco Demelas, Joseph Le Roux, Mathieu Lacroix 0001, Axel Parmentier |
ICML | 2 |
| 2022 | Exploiting Inductive Bias in Transformers for Unsupervised Disentanglement of Syntax and Semantics with VAEsabstractWe propose a generative model for text generation, which exhibits disentangled latent representations of syntax and semantics.Contrary to previous work, this model does not need syntactic information such as constituency parses, or semantic information such as paraphrase pairs.Our model relies solely on the inductive bias found in attention-based architectures such as Transformers.In the attention of Transformers, keys handle information selection while values specify what information is conveyed.Our model, dubbed QKVAE, uses Attention in its decoder to read latent variables where one latent variable infers keys while another infers values.We run experiments on latent representations and experiments on syntax/semantics transfer which show that QKVAE displays clear signs of disentangled syntax and semantics.We also show that our model displays competitive syntax transfer capabilities when compared to supervised models and that comparable supervised models need a fairly large amount of data (more than 50K samples) to outperform it on both syntactic and semantic transfer.The code for our experiments is publicly available 1 . Ghazi Felhi, Joseph Le Roux, Djamé Seddah |
NAACL-HLT | 2 |
| 2020 | Multitask Easy-First Dependency Parsing: Exploiting Complementarities of Different Dependency RepresentationsabstractWe present a parsing model for projective dependency trees which takes advantage of the existence of complementary dependency annotations for a language.This is the case for Arabic with the availability of CATiB and UD treebanks.Our system performs syntactic parsing according to both annotation types jointly as a sequence of arc-creating operations following the Easy-First approach, and partially created trees for one annotation type are also available to the other as features for the score function.This method gives error reduction of 9.9% on CATiB and 6.1% on UD compared to a single-task baseline, and ablation tests show that the main contribution of this reduction is given by sharing tree representation between tasks, and not simply sharing BiLSTM layers as is usually performed in NLP multitask systems. Yash Kankanampati, Joseph Le Roux, Nadi Tomeh, Dima Taji, Nizar Habash |
COLING | 2 |
| 2019 | Representation Learning and Dynamic Programming for Arc-Hybrid ParsingabstractInternational audience Joseph Le Roux, Antoine Rozenknop, Mathieu Lacroix 0001 |
CoNLL | 1 |
| 2017 | Efficient Discontinuous Phrase-Structure Parsing via the Generalized Maximum Spanning ArborescenceabstractWe present a new method for the joint task of tagging and non-projective dependency parsing.We demonstrate its usefulness with an application to discontinuous phrase-structure parsing where decoding lexicalized spines and syntactic derivations is performed jointly.The main contributions of this paper are (1) a reduction from joint tagging and non-projective dependency parsing to the Generalized Maximum Spanning Arborescence problem, and (2) a novel decoding algorithm for this problem through Lagrangian relaxation.We evaluate this model and obtain state-of-the-art results despite strong independence assumptions. Caio F. Corro, Joseph Le Roux, Mathieu Lacroix 0001 |
EMNLP | 2 |
| 2016 | Dependency Parsing with Bounded Block Degree and Well-nestedness via Lagrangian Relaxation and Branch-and-BoundabstractCaio Corro, Joseph Le Roux, Mathieu Lacroix, Antoine Rozenknop, Roberto Wolfler Calvo. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Caio F. Corro, Joseph Le Roux, Mathieu Lacroix 0001, Antoine Rozenknop, Roberto Wolfler Calvo |
ACL (1) | 2 |
| 2016 | Deep Lexical Segmentation and Syntactic Parsing in the Easy-First Dependency FrameworkabstractWe explore the consequences of representing token segmentations as hierarchical structures (trees) for the task of Multiword Expression (MWE) recognition, in isolation or in combination with dependency parsing.We propose a novel representation of token segmentation as trees on tokens, resembling dependency trees.Given this new representation, we present and evaluate two different architectures to combine MWE recognition and dependency parsing in the easy-first framework: a pipeline and a joint system, both taking advantage of lexical and syntactic dimensions.We experimentally validate that MWE recognition significantly helps syntactic parsing. Matthieu Constant, Joseph Le Roux, Nadi Tomeh |
HLT-NAACL | 2 |
| 2015 | Foreebank: Syntactic Analysis of Customer Support ForumsabstractWe present a new treebank of English and French technical forum content which has been annotated for grammatical errors and phrase structure.This double annotation allows us to empirically measure the effect of errors on parsing performance.While it is slightly easier to parse the corrected versions of the forum sentences, the errors are not the main factor in making this kind of text hard to parse. Rasoul Samad Zadeh Kaljahi, Jennifer Foster, Johann Roturier, Corentin Ribeyre, Teresa Lynn, Joseph Le Roux |
EMNLP | 6 |
| 2014 | Syntactic Parsing and Compound Recognition via Dual Decomposition: Application to French
Joseph Le Roux, Antoine Rozenknop, Matthieu Constant |
COLING | 1 |
| 2013 | Combining PCFG-LA Models with Dual Decomposition: A Case Study with Function Labels and BinarizationabstractIt has recently been shown that different NLP models can be effectively combined using dual decomposition.In this paper we demonstrate that PCFG-LA parsing models are suitable for combination in this way.We experiment with the different models which result from alternative methods of extracting a grammar from a treebank (retaining or discarding function labels, left binarization versus right binarization) and achieve a labeled Parseval F-score of 92.4 on Wall Street Journal Section 23 -this represents an absolute improvement of 0.7 and an error reduction rate of 7% over a strong PCFG-LA product-model baseline.Although we experiment only with binarization and function labels in this study, there is much scope for applying this approach to other grammar extraction strategies. Joseph Le Roux, Antoine Rozenknop, Jennifer Foster |
EMNLP | 1 |
| 2013 | XMG: eXtensible MetaGrammarabstractIn this article, we introduce eXtensible MetaGrammar (XMG), a framework for specifying tree-based grammars such as Feature-Based Lexicalized Tree-Adjoining Grammars (FB-LTAG) and Interaction Grammars (IG). We argue that XMG displays three features that facilitate both grammar writing and a fast prototyping of tree-based grammars. Firstly, XMG is fully declarative. For instance, it permits a declarative treatment of diathesis that markedly departs from the procedural lexical rules often used to specify tree-based grammars. Secondly, the XMG language has a high notational expressivity in that it supports multiple linguistic dimensions, inheritance, and a sophisticated treatment of identifiers. Thirdly, XMG is extensible in that its computational architecture facilitates the extension to other linguistic formalisms. We explain how this architecture naturally supports the design of three linguistic formalisms, namely, FB-LTAG, IG, and Multi-Component Tree-Adjoining Grammar (MC-TAG). We further show how it permits a straightforward integration of additional mechanisms such as linguistic and formal principles. To further illustrate the declarativity, notational expressivity, and extensibility of XMG, we describe the methodology used to specify an FB-LTAG for French augmented with a unification-based compositional semantics. This illustrates both how XMG facilitates the modeling of the tree fragment hierarchies required to specify tree-based grammars and of a syntax/semantics interface between semantic representations and syntactic trees. Finally, we briefly report on several grammars for French, English, and German that were implemented using XMG and compare XMG with other existing grammar specification frameworks for tree-based grammars. Benoît Crabbé, Denys Duchier, Claire Gardent, Joseph Le Roux, Yannick Parmentier 0001 |
Comput. Linguistics | 4 |
| 2012 | Semi-supervised Dependency Parsing using Lexical Affinities
Seyed Abolghasem Mirroshandel, Alexis Nasr, Joseph Le Roux |
ACL (1) | 3 |
| 2011 | From News to Comment: Resources and Benchmarks for Parsing the Language of Web 2.0
Jennifer Foster, Özlem Çetinoglu, Joachim Wagner 0001, Joseph Le Roux, Joakim Nivre, Deirdre Hogan, Josef van Genabith |
IJCNLP | 4 |
| 2006 | XMG - An Expressive Formalism for Describing Tree-Based Grammars
Yannick Parmentier 0001, Joseph Le Roux, Benoît Crabbé |
EACL | 2 |
| 2006 | Lexical Disambiguation with Polarities and Automata
Guillaume Bonfante, Joseph Le Roux, Guy Perrier |
CIAA | 2 |