VLDB 2026 Research / reviewers in the wild / expert
C. Guetta
dblp:438/1973
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Optimization for machine learning · 29% Probabilistic and Bayesian machine learning · 14% Deep learning architectures and training · 14% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
0.9 | 1 | 2025 | Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › pre-training
data mixture optimization |
0.9 | 1 | 2025 | Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining |
0.9 | 1 | 2025 | Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework · NeurIPS 2025 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
multi-fidelity bayesian optimization |
0.9 | 1 | 2025 | Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.9 | 1 | 2025 | Architectural and Inferential Inductive Biases for Exchangeable Sequence Modeling · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.9 | 1 | 2025 | Architectural and Inferential Inductive Biases for Exchangeable Sequence Modeling · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
scaling laws · 0.9probabilistic extrapolation · 0.9posterior inference · 0.9multi-fidelity bayesian optimization · 0.9causal masking · 0.9autoregressive generation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Architectural and Inferential Inductive Biases for Exchangeable Sequence ModelingabstractAutoregressive models have emerged as a powerful framework for modeling exchangeable sequences---i.i.d. observations when conditioned on some latent factor---enabling direct modeling of uncertainty from missing data (rather than a latent). Motivated by the critical role posterior inference plays as a subroutine in decision-making (e.g., active learning, bandits), we study the inferential and architectural inductive biases that are most effective for exchangeable sequence modeling. For the inference stage, we highlight a fundamental limitation of the prevalent single-step generation approach: its inability to distinguish between epistemic and aleatoric uncertainty. Instead, a long line of works in Bayesian statistics advocates for multi-step autoregressive generation; we demonstrate this "correct approach" enables superior uncertainty quantification that translates into better performance on downstream decision-making tasks. This naturally leads to the next question: which architectures are best suited for multi-step inference? We identify a subtle yet important gap between recently proposed Transformer architectures for exchangeable sequences (Müller et al., 2022; Nguyen & Grover, 2022; Ye & Namkoong, 2024), and prove that they in fact cannot guarantee exchangeability despite introducing significant computational overhead. Through empirical evaluation, we find that these custom architectures can significantly underperform compared to standard causal masking, highlighting the need for new architectural innovations in Transformer-based modeling of exchangeable sequences. Daksh Mittal, Leon Li, Thomson Yen, C. Guetta, Hongseok Namkoong |
NeurIPS | 4 |
| 2025 | Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian FrameworkabstractCareful curation of data sources can significantly improve the performance of LLM pre-training, but predominant approaches rely heavily on intuition or costly trial-and-error, making them difficult to generalize across different data domains and downstream tasks. Although scaling laws can provide a principled and general approach for data curation, standard deterministic extrapolation from small-scale experiments to larger scales requires strong assumptions on the reliability of such extrapolation, whose brittleness has been highlighted in prior works. In this paper, we introduce a probabilistic extrapolation framework for data mixture optimization that avoids rigid assumptions and explicitly models the uncertainty in performance across decision variables. We formulate data curation as a sequential decision-making problem–multi-fidelity, multi-scale Bayesian optimization–where {data mixtures, model scale, training steps} are adaptively selected to balance training cost and potential information gain. Our framework naturally gives rise to algorithm prototypes that leverage noisy information from inexpensive experiments to systematically inform costly training decisions. To accelerate methodological progress, we build a simulator based on 472 language model pre-training runs with varying data compositions from the SlimPajama dataset. We observe that even simple kernels and acquisition functions can enable principled decisions across training models from 20M to 1B parameters and achieve 2.6x and 3.3x speedups compared to multi-fidelity BO and random search baselines. Taken together, our framework underscores potential efficiency gains achievable by developing principled and transferable data mixture optimization methods. Our code is publicly available at https://github.com/namkoong-lab/data-recipes. Thomson Yen, Andrew Wei Tung Siah, C. Guetta, Tianyi Peng, Hongseok Namkoong |
NeurIPS | 4 |