VLDB 2026 Research / reviewers in the wild / expert
Jean Kaddour
dblp:232/9592
· DBLP profile ↗
9ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 8 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Probabilistic and Bayesian machine learning · 23% Reinforcement learning · 20% Language models and text generation · 14% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
procedural task generation |
0.9 | 1 | 2025 | Reasoning Gym: Reasoning Environments for Reinforcement Learning with Verifiable Rewards · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model
reasoning model |
0.9 | 1 | 2025 | Reasoning Gym: Reasoning Environments for Reinforcement Learning with Verifiable Rewards · NeurIPS 2025 |
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards |
0.9 | 1 | 2025 | Reasoning Gym: Reasoning Environments for Reinforcement Learning with Verifiable Rewards · NeurIPS 2025 |
Program synthesis and code generation › code generation evaluation
code generation benchmark |
0.9 | 1 | 2025 | BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions · ICLR 2025 |
Program synthesis and code generation
code generation with language models |
0.9 | 1 | 2025 | BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions · ICLR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery |
0.7 | 1 | 2023 | DAG Learning on the Permutahedron · ICLR 2023 |
Machine learning › Optimization for machine learning
combinatorial optimization |
0.7 | 1 | 2023 | DAG Learning on the Permutahedron · ICLR 2023 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
directed acyclic graph learning |
0.7 | 1 | 2023 | DAG Learning on the Permutahedron · ICLR 2023 |
Machine learning › Efficient and distributed learning
efficient training |
0.7 | 1 | 2023 | No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models · NeurIPS 2023 |
Machine learning › Graph learning
graph self-supervised learning |
0.7 | 1 | 2023 | Evaluating Self-Supervised Learning for Molecular Graph Embeddings · NeurIPS 2023 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
transformer language model |
0.7 | 1 | 2023 | No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models · NeurIPS 2023 |
Machine learning › Learning theory
generalization |
0.6 | 1 | 2022 | When Do Flat Minima Optimizers Work? · NeurIPS 2022 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.5 | 1 | 2021 | Causal Effect Inference for Structured Treatments · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › heterogeneous treatment effect estimation
conditional average treatment effect |
0.5 | 1 | 2021 | Causal Effect Inference for Structured Treatments · NeurIPS 2021 |
Machine learning › Efficient and distributed learning
active learning |
0.4 | 1 | 2020 | Probabilistic Active Meta-Learning · NeurIPS 2020 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.4 | 1 | 2020 | Probabilistic Active Meta-Learning · NeurIPS 2020 |
Machine learning › Reinforcement learning › curriculum reinforcement learning
task selection |
0.4 | 1 | 2020 | Probabilistic Active Meta-Learning · NeurIPS 2020 |
Machine learning › Representation and self-supervised learning
pre-training |
0.2 | 1 | 2023 | No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models · NeurIPS 2023 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.2 | 1 | 2022 | When Do Flat Minima Optimizers Work? · NeurIPS 2022 |
Machine learning › Graph learning
graph neural network |
0.1 | 1 | 2021 | Causal Effect Inference for Structured Treatments · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.1 | 1 | 2020 | Probabilistic Active Meta-Learning · NeurIPS 2020 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.7procedural generation · 1.7large language model · 0.9benchmark construction · 0.9self-supervised learning · 0.7permutahedron optimization · 0.7graph neural network · 0.7efficient optimizers · 0.7dynamic architecture · 0.7batch selection · 0.7stochastic weight averaging · 0.6sharpness-aware minimization · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsabstractTask automation has been greatly empowered by the recent advances in Large Language Models (LLMs) via Python code, where the tasks range from software engineering development to general-purpose reasoning. While current benchmarks have shown that LLMs can solve tasks using programs like human developers, the majority of their evaluations are limited to short and self-contained algorithmic tasks or standalone function calls. Solving challenging and practical tasks requires the capability of utilizing **diverse function calls as tools** to efficiently implement functionalities like data analysis and web development. In addition, using multiple tools to solve a task needs compositional reasoning by accurately understanding **complex instructions**. Fulfilling both of these characteristics can pose a great challenge for LLMs. To assess how well LLMs can solve challenging and practical tasks via programs, we introduce BigCodeBench, a benchmark that challenges LLMs to invoke multiple function calls as tools from 139 libraries and 7 domains for 1,140 fine-grained tasks. To evaluate LLMs rigorously, each task encompasses 5.6 test cases with an average branch coverage of 99%. In addition, we propose a natural-language-oriented variant of BigCodeBench, BigCodeBench-Instruct, that automatically transforms the original docstrings into short instructions containing only essential information. Our extensive evaluation of 60 LLMs shows that **LLMs are not yet capable of following complex instructions to use function calls precisely, with scores up to 60%, significantly lower than the human performance of 97%**. The results underscore the need for further advancements in this area. Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu 0011, Wenhao Yu 0002, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, Simon Brunner, Chen Gong 0005, James Hoang, Armel Zebaze, Xiaoheng Hong, Wen-Ding Li, Jean Kaddour, Zhihan Zhang 0001, Prateek Yadav |
ICLR | 17 |
| 2025 | Are We Done with MMLU?abstractAryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile Van Krieken, Pasquale Minervini. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao 0043, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile van Krieken, Pasquale Minervini |
NAACL (Long Papers) | 14 |
| 2025 | Reasoning Gym: Reasoning Environments for Reinforcement Learning with Verifiable RewardsabstractWe introduce Reasoning Gym, a library of reasoning environments for reinforcement learning with verifiable rewards (RLVR). It provides over 100 tasks spanning multiple domains including algebra, arithmetic, computation, cognition, geometry, graph theory, logic, and various common games. Its key innovation is the ability to generate virtually infinite training data with adjustable complexity, unlike most previous reasoning datasets, which are typically fixed. This procedural generation approach allows for continuous evaluation across varying difficulty levels and task configurations. Our experimental results demonstrate the efficacy of Reasoning Gym in both evaluating and reinforcement learning of reasoning models. Zafir Stojanovski, Oliver Stanley, Joe Sharratt, Abdulhakeem Adefioye, Jean Kaddour, Andreas Köpf |
NeurIPS | 6 |
| 2023 | DAG Learning on the Permutahedron
Valentina Zantedeschi, Luca Franceschi 0001, Jean Kaddour, Matt J. Kusner, Vlad Niculae |
ICLR | 3 |
| 2023 | No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language ModelsabstractThe computation necessary for training Transformer-based language models has skyrocketed in recent years.
This trend has motivated research on efficient training algorithms designed to improve training, validation, and downstream performance faster than standard training. In this work, we revisit three categories of such algorithms: dynamic architectures (layer stacking, layer dropping), batch selection (selective backprop., RHO-loss), and efficient optimizers (Lion, Sophia). When pre-training BERT and T5 with a fixed computation budget using such methods, we find that their training, validation, and downstream gains vanish compared to a baseline with a fully-decayed learning rate. We define an evaluation protocol that enables computation to be done on arbitrary machines by mapping all computation time to a reference machine which we call reference system time. We discuss the limitations of our proposed protocol and release our code to encourage rigorous research in efficient training procedures: https://github.com/JeanKaddour/NoTrainNoGain. Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, Matt J. Kusner |
NeurIPS | 1 |
| 2023 | Evaluating Self-Supervised Learning for Molecular Graph EmbeddingsabstractGraph Self-Supervised Learning (GSSL) provides a robust pathway for acquiring embeddings without expert labelling, a capability that carries profound implications for molecular graphs due to the staggering number of potential molecules and the high cost of obtaining labels. However, GSSL methods are designed not for optimisation within a specific domain but rather for transferability across a variety of downstream tasks. This broad applicability complicates their evaluation. Addressing this challenge, we present "Molecular Graph Representation Evaluation" (MOLGRAPHEVAL), generating detailed profiles of molecular graph embeddings with interpretable and diversified attributes. MOLGRAPHEVAL offers a suite of probing tasks grouped into three categories: (i) generic graph, (ii) molecular substructure, and (iii) embedding space properties. By leveraging MOLGRAPHEVAL to benchmark existing GSSL methods against both current downstream datasets and our suite of tasks, we uncover significant inconsistencies between inferences drawn solely from existing datasets and those derived from more nuanced probing. These findings suggest that current evaluation methodologies fail to capture the entirety of the landscape. Hanchen Wang 0002, Jean Kaddour, Shengchao Liu, Jian Tang 0005, Joan Lasenby, Qi Liu 0049 |
NeurIPS | 2 |
| 2022 | When Do Flat Minima Optimizers Work?abstractRecently, flat-minima optimizers, which seek to find parameters in low-loss neighborhoods, have been shown to improve a neural network's generalization performance over stochastic and adaptive gradient-based optimizers. Two methods have received significant attention due to their scalability: 1. Stochastic Weight Averaging (SWA), and 2. Sharpness-Aware Minimization (SAM). However, there has been limited investigation into their properties and no systematic benchmarking of them across different domains. We fill this gap here by comparing the loss surfaces of the models trained with each method and through broad benchmarking across computer vision, natural language processing, and graph representation learning tasks. We discover several surprising findings from these results, which we hope will help researchers further improve deep learning optimizers, and practitioners identify the right optimizer for their problem. Jean Kaddour, Linqing Liu, Ricardo Silva 0001, Matt J. Kusner |
NeurIPS | 1 |
| 2021 | Causal Effect Inference for Structured TreatmentsabstractWe address the estimation of conditional average treatment effects (CATEs) for structured treatments (e.g., graphs, images, texts). Given a weak condition on the effect, we propose the generalized Robinson decomposition, which (i) isolates the causal estimand (reducing regularization bias), (ii) allows one to plug in arbitrary models for learning, and (iii) possesses a quasi-oracle convergence guarantee under mild assumptions. In experiments with small-world and molecular graphs we demonstrate that our approach outperforms prior work in CATE estimation. Jean Kaddour, Qi Liu 0049, Matt J. Kusner, Ricardo Silva 0001 |
NeurIPS | 1 |
| 2020 | Probabilistic Active Meta-LearningabstractData-efficient learning algorithms are essential in many practical applications where data collection is expensive, e.g., in robotics due to the wear and tear. To address this problem, meta-learning algorithms use prior experience about tasks to learn new, related tasks efficiently. Typically, a set of training tasks is assumed given or randomly chosen. However, this setting does not take into account the sequential nature that naturally arises when training a model from scratch in real-life: how do we collect a set of training tasks in a data-efficient manner? In this work, we introduce task selection based on prior experience into a meta-learning algorithm by conceptualizing the learner and the active meta-learning setting using a probabilistic latent variable model. We provide empirical evidence that our approach improves data-efficiency when compared to strong baselines on simulated robotic experiments. Jean Kaddour, Steindór Sæmundsson, Marc Peter Deisenroth |
NeurIPS | 1 |