VLDB 2026 Research / reviewers in the wild / expert
Léo Grinsztajn
dblp:259/3203
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Deep learning architectures and training · 50% Transfer learning and domain adaptation · 18% Kernel, tree and ensemble methods · 18% | |
| Databases, data mining, and information retrieval
2 papers |
Machine learning and data management · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Kernel, tree and ensemble methods › gradient boosting
gradient boosted decision trees |
0.8 | 1 | 2024 | Better by default: Strong pre-tuned MLPs and boosted trees on tabular data · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › feedforward neural network
multilayer perceptron |
0.8 | 1 | 2024 | Better by default: Strong pre-tuned MLPs and boosted trees on tabular data · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
tabular data learning |
0.8 | 1 | 2024 | Better by default: Strong pre-tuned MLPs and boosted trees on tabular data · NeurIPS 2024 |
Machine learning and data management › table representation learning
table pre-training |
0.8 | 1 | 2024 | CARTE: Pretraining and Transfer for Tabular Learning · ICML 2024 |
Machine learning › Learning theory
inductive bias |
0.6 | 1 | 2022 | Why do tree-based models still outperform deep learning on typical tabular data? · NeurIPS 2022 |
Machine learning › Deep learning architectures and training
tabular deep learning |
0.6 | 1 | 2022 | Why do tree-based models still outperform deep learning on typical tabular data? · NeurIPS 2022 |
Machine learning and data management
tabular data benchmark |
0.2 | 1 | 2022 | Why do tree-based models still outperform deep learning on typical tabular data? · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
string embedding · 1.5graph representation · 1.5graph attention network · 1.5random forest · 1.1hyperparameter search · 1.1gradient boosting · 1.1meta-tuning · 0.8hyperparameter optimization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CARTE: Pretraining and Transfer for Tabular LearningabstractPretrained deep-learning models are the go-to solution for images or text. However, for tabular data the standard is still to train tree-based models. Indeed, transfer learning on tables hits the challenge of data integration: finding correspondences, correspondences in the entries (entity matching) where different words may denote the same entity, correspondences across columns (schema matching), which may come in different orders, names... We propose a neural architecture that does not need such correspondences. As a result, we can pretrain it on background data that has not been matched. The architecture –CARTE for Context Aware Representation of Table Entries– uses a graph representation of tabular (or relational) data to process tables with different columns, string embedding of entries and columns names to model an open vocabulary, and a graph-attentional network to contextualize entries with column names and neighboring entries. An extensive benchmark shows that CARTE facilitates learning, outperforming a solid set of baselines including the best tree-based models. CARTE also enables joint learning across tables with unmatched columns, enhancing a small table with bigger ones. CARTE opens the door to large pretrained models for tabular data. Myung Jun Kim, Léo Grinsztajn, Gaël Varoquaux |
ICML | 2 |
| 2024 | Better by default: Strong pre-tuned MLPs and boosted trees on tabular dataabstractFor classification and regression on tabular data, the dominance of gradient-boosted decision trees (GBDTs) has recently been challenged by often much slower deep learning methods with extensive hyperparameter tuning. We address this discrepancy by introducing (a) RealMLP, an improved multilayer perceptron (MLP), and (b) strong meta-tuned default parameters for GBDTs and RealMLP. We tune RealMLP and the default parameters on a meta-train benchmark with 118 datasets and compare them to hyperparameter-optimized versions on a disjoint meta-test benchmark with 90 datasets, as well as the GBDT-friendly benchmark by Grinsztajn et al. (2022). Our benchmark results on medium-to-large tabular datasets (1K--500K samples) show that RealMLP offers a favorable time-accuracy tradeoff compared to other neural baselines and is competitive with GBDTs in terms of benchmark scores. Moreover, a combination of RealMLP and GBDTs with improved default parameters can achieve excellent results without hyperparameter tuning. Finally, we demonstrate that some of RealMLP's improvements can also considerably improve the performance of TabR with default parameters. David Holzmüller, Léo Grinsztajn, Ingo Steinwart |
NeurIPS | 2 |
| 2022 | Why do tree-based models still outperform deep learning on typical tabular data?abstractWhile deep learning has enabled tremendous progress on text and image datasets, its superiority on tabular data is not clear. We contribute extensive benchmarks of standard and novel deep learning methods as well as tree-based models such as XGBoost and Random Forests, across a large number of datasets and hyperparameter combinations. We define a standard set of 45 datasets from varied domains with clear characteristics of tabular data and a benchmarking methodology accounting for both fitting models and finding good hyperparameters. Results show that tree-based models remain state-of-the-art on medium-sized data ($\sim$10K samples) even without accounting for their superior speed. To understand this gap, we conduct an empirical investigation into the differing inductive biases of tree-based models and neural networks. This leads to a series of challenges which should guide researchers aiming to build tabular-specific neural network: 1) be robust to uninformative features, 2) preserve the orientation of the data, and 3) be able to easily learn irregular functions. To stimulate research on tabular architectures, we contribute a standard benchmark and raw data for baselines: every point of a 20\,000 compute hours hyperparameter search for each learner. Léo Grinsztajn, Edouard Oyallon, Gaël Varoquaux |
NeurIPS | 1 |