VLDB 2026 Research / reviewers in the wild / expert
Nicolas Huynh
dblp:134/9604
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Generative modeling · 49% Optimization for machine learning · 13% Trustworthy machine learning · 13% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Kernel, tree and ensemble methods
decision tree |
0.9 | 1 | 2025 | Decision Tree Induction Through LLMs via Semantically-Aware Evolution · ICLR 2025 |
Machine learning › Optimization for machine learning
evolutionary computation |
0.9 | 1 | 2025 | Decision Tree Induction Through LLMs via Semantically-Aware Evolution · ICLR 2025 |
Machine learning › Trustworthy machine learning › interpretability
explainable AI |
0.9 | 1 | 2025 | Decision Tree Induction Through LLMs via Semantically-Aware Evolution · ICLR 2025 |
Machine learning › Deep learning architectures and training
data augmentation |
0.8 | 1 | 2024 | Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes · ICML 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Time Series Diffusion in the Frequency Domain · ICML 2024 |
Machine learning › Generative modeling › synthetic data generation
LLM-based data generation |
0.8 | 1 | 2024 | Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes · ICML 2024 |
Machine learning › Generative modeling › synthetic data generation
tabular data augmentation |
0.8 | 1 | 2024 | Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes · ICML 2024 |
Machine learning › Generative modeling › diffusion model › time series generation
time series diffusion models |
0.8 | 1 | 2024 | Time Series Diffusion in the Frequency Domain · ICML 2024 |
Machine learning › Generative modeling
score-based model |
0.2 | 1 | 2024 | Time Series Diffusion in the Frequency Domain · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
large language model · 1.6semantic priors · 0.9genetic programming · 0.9fitness-guided crossover · 0.9diversity-guided mutation · 0.9uncertainty metrics · 0.8stochastic differential equation · 0.8fourier analysis · 0.8denoising score matching · 0.8data curation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Decision Tree Induction Through LLMs via Semantically-Aware EvolutionabstractDecision trees are a crucial class of models offering robust predictive performance and inherent interpretability across various domains, including healthcare, finance, and logistics. However, current tree induction methods often face limitations such as suboptimal solutions from greedy methods or prohibitive computational costs and limited applicability of exact optimization approaches.
To address these challenges, we propose an evolutionary optimization method for decision tree induction based on genetic programming (GP). Our key innovation is the integration of semantic priors and domain-specific knowledge about the search space into the optimization algorithm.
To this end, we introduce $\texttt{LLEGO}$, a framework that incorporates semantic priors into genetic search operators through the use of Large Language Models (LLMs), thereby enhancing search efficiency and targeting regions of the search space that yield decision trees with superior generalization performance. This is operationalized through novel genetic operators that work with structured natural language prompts, effectively utilizing LLMs as conditional generative models and sources of semantic knowledge. Specifically, we introduce $\textit{fitness-guided}$ crossover to exploit high-performing regions, and $\textit{diversity-guided}$ mutation for efficient global exploration of the search space. These operators are controlled by corresponding hyperparameters that enable a more nuanced balance between exploration and exploitation across the search space. Empirically, we demonstrate across various benchmarks that $\texttt{LLEGO}$ evolves superior-performing trees compared to existing tree induction methods, and exhibits significantly more efficient search performance compared to conventional GP approaches. Tennison Liu, Nicolas Huynh, Mihaela van der Schaar |
ICLR | 2 |
| 2024 | DAGnosis: Localized Identification of Data Inconsistencies using StructuresabstractIdentification and appropriate handling of inconsistencies in data at deployment time is crucial to reliably use machine learning models. While recent data-centric methods are able to identify such inconsistencies with respect to the training set, they suffer from two key limitations: (1) suboptimality in settings where features exhibit statistical independencies, due to their usage of compressive representations and (2) lack of localization to pin-point why a sample might be flagged as inconsistent, which is important to guide future data collection. We solve these two fundamental limitations using directed acyclic graphs (DAGs) to encode the training set’s features probability distribution and independencies as a structure. Our method, called DAGnosis, leverages these structural interactions to bring valuable and insightful data-centric conclusions. DAGnosis unlocks the localization of the causes of inconsistencies on a DAG, an aspect overlooked by previous approaches. Moreover, we show empirically that leveraging these interactions (1) leads to more accurate conclusions in detecting inconsistencies, as well as (2) provides more detailed insights into why some samples are flagged. Nicolas Huynh, Jeroen Berrevoets, Nabeel Seedat, Jonathan Crabbé, Zhaozhi Qian, Mihaela van der Schaar |
AISTATS | 1 |
| 2024 | Time Series Diffusion in the Frequency DomainabstractFourier analysis has been an instrumental tool in the development of signal processing. This leads us to wonder whether this framework could similarly benefit generative modelling. In this paper, we explore this question through the scope of time series diffusion models. More specifically, we analyze whether representing time series in the frequency domain is a useful inductive bias for score-based diffusion models. By starting from the canonical SDE formulation of diffusion in the time domain, we show that a dual diffusion process occurs in the frequency domain with an important nuance: Brownian motions are replaced by what we call mirrored Brownian motions, characterized by mirror symmetries among their components. Building on this insight, we show how to adapt the denoising score matching approach to implement diffusion models in the frequency domain. This results in frequency diffusion models, which we compare to canonical time diffusion models. Our empirical evaluation on real-world datasets, covering various domains like healthcare and finance, shows that frequency diffusion models better capture the training distribution than time diffusion models. We explain this observation by showing that time series from these datasets tend to be more localized in the frequency domain than in the time domain, which makes them easier to model in the former case. All our observations point towards impactful synergies between Fourier analysis and diffusion models. Jonathan Crabbé, Nicolas Huynh, Jan Stanczuk, Mihaela van der Schaar |
ICML | 2 |
| 2024 | Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimesabstractMachine Learning (ML) in low-data settings remains an underappreciated yet crucial problem. Hence, data augmentation methods to increase the sample size of datasets needed for ML are key to unlocking the transformative potential of ML in data-deprived regions and domains. Unfortunately, the limited training set constrains traditional tabular synthetic data generators in their ability to generate a large and diverse augmented dataset needed for ML tasks. To address this challenge, we introduce $\texttt{CLLM}$, which leverages the prior knowledge of Large Language Models (LLMs) for data augmentation in the low-data regime. However, not all the data generated by LLMs will improve downstream utility, as for any generative model. Consequently, we introduce a principled curation mechanism, leveraging learning dynamics, coupled with confidence and uncertainty metrics, to obtain a high-quality dataset. Empirically, on multiple real-world datasets, we demonstrate the superior performance of $\texttt{CLLM}$ in the low-data regime compared to conventional generators. Additionally, we provide insights into the LLM generation and curation mechanism, shedding light on the features that enable them to output high-quality augmented datasets. Nabeel Seedat, Nicolas Huynh, Boris van Breugel, Mihaela van der Schaar |
ICML | 2 |