VLDB 2026 Research / reviewers in the wild / expert
Daniel Augusto R. M. A. de Souza
dblp:244/1958 · also Daniel Augusto de Souza
· DBLP profile ↗
10ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-4721-2401ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Probabilistic and Bayesian machine learning · 37% Deep learning architectures and training · 17% Representation and self-supervised learning · 13% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
1.5 | 2 | 2025 | Infinite Neural Operators: Gaussian processes on functions · NeurIPS 2025 Thin and deep Gaussian processes · NeurIPS 2023 |
Machine learning › Deep learning architectures and training › neural operator
fourier neural operator |
0.9 | 1 | 2025 | Infinite Neural Operators: Gaussian processes on functions · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
neural operator |
0.9 | 1 | 2025 | Infinite Neural Operators: Gaussian processes on functions · NeurIPS 2025 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning |
0.9 | 1 | 2025 | When Can Proxies Improve the Sample Complexity of Preference Learning? · ICML 2025 |
Natural language and speech › Language models and text generation › alignment
reward hacking |
0.9 | 1 | 2025 | When Can Proxies Improve the Sample Complexity of Preference Learning? · ICML 2025 |
Machine learning › Learning theory
sample complexity |
0.9 | 1 | 2025 | When Can Proxies Improve the Sample Complexity of Preference Learning? · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.8 | 1 | 2024 | Streaming Bayes GFlowNets · NeurIPS 2024 |
Machine learning › Generative modeling
generative flow networks |
0.8 | 1 | 2024 | Streaming Bayes GFlowNets · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
online inference |
0.8 | 1 | 2024 | Streaming Bayes GFlowNets · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › hierarchical gaussian process
deep gaussian process |
0.7 | 1 | 2023 | Thin and deep Gaussian processes · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning
latent representation |
0.7 | 1 | 2023 | Thin and deep Gaussian processes · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning
low-dimensional manifold |
0.7 | 1 | 2023 | Thin and deep Gaussian processes · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
reward regularization · 0.9reward model · 0.9kernel methods · 0.9gaussian process · 0.9variational inference · 0.8amortized sampling · 0.8lengthscale parameterization · 0.7kernel parameterization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting kernel complexity in Gaussian process regression: An empirical study on text-to-visual embedding mapping in fashion design
Dongmei Mo, Daniel Augusto R. M. A. de Souza, Xingxing Zou, Wai Keung Wong |
Neurocomputing | 2 |
| 2025 | When Can Proxies Improve the Sample Complexity of Preference Learning?abstractWe address the problem of reward hacking, where maximising a proxy reward does not necessarily increase the true reward. This is a key concern for Large Language Models (LLMs), as they are often fine-tuned on human preferences that may not accurately reflect a true objective. Existing work uses various tricks such as regularisation, tweaks to the reward model, and reward hacking detectors, to limit the influence that such proxy preferences have on a model. Luckily, in many contexts such as medicine, education, and law, a sparse amount of expert data is often available. In these cases, it is often unclear whether the addition of proxy data can improve policy learning. We outline a set of sufficient conditions on proxy feedback that, if satisfied, indicate that proxy data can provably improve the sample complexity of learning the ground truth policy. These conditions can inform the data collection process for specific tasks. The result implies a parameterisation for LLMs that achieves this improved sample complexity. We detail how one can adapt existing architectures to yield this improved sample complexity. Daniel Augusto R. M. A. de Souza, Zhengyan Shi, Mengyue Yang, Pasquale Minervini, Matt J. Kusner, Alexander D'Amour |
ICML | 2 |
| 2025 | Infinite Neural Operators: Gaussian processes on functionsabstractA variety of infinitely wide neural architectures (e.g., dense NNs, CNNs, and transformers) induce Gaussian process (GP) priors over their outputs.
These relationships provide both an accurate characterization of the prior predictive distribution and enable the use of GP machinery to improve the uncertainty quantification of deep neural networks.
In this work, we extend this connection to neural operators (NOs), a class of models designed to learn mappings between function spaces.
Specifically, we show conditions for when arbitrary-depth NOs with Gaussian-distributed convolution kernels converge to function-valued GPs.
Based on this result, we show how to compute the covariance functions of these NO-GPs for two NO parametrizations, including the popular Fourier neural operator (FNO).
With this, we compute the posteriors of these GPs in regression scenarios, including PDE solution operators.
This work is an important step towards uncovering the inductive biases of current FNO architectures and opens a path to incorporate novel inductive biases for use in kernel-based operator learning methods. Daniel Augusto R. M. A. de Souza, Jake Cunningham, Yuri F. Saporito, Diego Mesquita, Marc Peter Deisenroth |
NeurIPS | 1 |
| 2024 | Streaming Bayes GFlowNetsabstractBayes' rule naturally allows for inference refinement in a streaming fashion, without the need to recompute posteriors from scratch whenever new data arrives. In principle, Bayesian streaming is straightforward: we update our prior with the available data and use the resulting posterior as a prior when processing the next data chunk. In practice, however, this recipe entails i) approximating an intractable posterior at each time step; and ii) encapsulating results appropriately to allow for posterior propagation. For continuous state spaces, variational inference (VI) is particularly convenient due to its scalability and the tractability of variational posteriors, For discrete state spaces, however, state-of-the-art VI results in analytically intractable approximations that are ill-suited for streaming settings. To enable streaming Bayesian inference over discrete parameter spaces, we propose streaming Bayes GFlowNets (abbreviated as SB-GFlowNets) by leveraging the recently proposed GFlowNets --- a powerful class of amortized samplers for discrete compositional objects. Notably, SB-GFlowNet approximates the initial posterior using a standard GFlowNet and subsequently updates it using a tailored procedure that requires only the newly observed data. Our case studies in linear preference learning and phylogenetic inference showcase the effectiveness of SB-GFlowNets in sampling from an unnormalized posterior in a streaming setting. As expected, we also observe that SB-GFlowNets is significantly faster than repeatedly training a GFlowNet from scratch to sample from the full posterior. Tiago da Silva, Daniel Augusto R. M. A. de Souza, Diego Mesquita |
NeurIPS | 2 |
| 2023 | Actually Sparse Variational Gaussian ProcessesabstractGaussian processes (GPs) are typically criticised for their unfavourable scaling in both computational and memory requirements. For large datasets, sparse GPs reduce these demands by conditioning on a small set of inducing variables designed to summarise the data. In practice however, for large datasets requiring many inducing variables, such as low-lengthscale spatial data, even sparse GPs can become computationally expensive, limited by the number of inducing variables one can use. In this work, we propose a new class of inter-domain variational GP, constructed by projecting a GP onto a set of compactly supported B-spline basis functions. The key benefit of our approach is that the compact support of the B-spline basis functions admits the use of sparse linear algebra to significantly speed up matrix operations and drastically reduce the memory footprint. This allows us to very efficiently model fast-varying spatial phenomena with tens of thousands of inducing variables, where previous approaches failed. Jake Cunningham, Daniel Augusto R. M. A. de Souza, So Takao, Mark van der Wilk, Marc Peter Deisenroth |
AISTATS | 2 |
| 2023 | Thin and deep Gaussian processesabstractGaussian processes (GPs) can provide a principled approach to uncertainty quantification with easy-to-interpret kernel hyperparameters, such as the lengthscale, which controls the correlation distance of function values.However, selecting an appropriate kernel can be challenging.
Deep GPs avoid manual kernel engineering by successively parameterizing kernels with GP layers, allowing them to learn low-dimensional embeddings of the inputs that explain the output data.
Following the architecture of deep neural networks, the most common deep GPs warp the input space layer-by-layer but lose all the interpretability of shallow GPs. An alternative construction is to successively parameterize the lengthscale of a kernel, improving the interpretability but ultimately giving away the notion of learning lower-dimensional embeddings. Unfortunately, both methods are susceptible to particular pathologies which may hinder fitting and limit their interpretability.
This work proposes a novel synthesis of both previous approaches: {Thin and Deep GP} (TDGP). Each TDGP layer defines locally linear transformations of the original input data maintaining the concept of latent embeddings while also retaining the interpretation of lengthscales of a kernel. Moreover, unlike the prior solutions, TDGP induces non-pathological manifolds that admit learning lower-dimensional representations.
We show with theoretical and experimental results that i) TDGP is, unlike previous models, tailored to specifically discover lower-dimensional manifolds in the input data, ii) TDGP behaves well when increasing the number of layers, and iii) TDGP performs well in standard benchmark datasets. Daniel Augusto R. M. A. de Souza, Alexander Nikitin 0002, St John, Magnus Ross, Mauricio A. Álvarez, Marc Peter Deisenroth, João Paulo Pordeus Gomes, Diego Mesquita, César Lincoln C. Mattos |
NeurIPS | 1 |
| 2022 | Parallel MCMC Without Embarrassing FailuresabstractEmbarrassingly parallel Markov Chain Monte Carlo (MCMC) exploits parallel computing to scale Bayesian inference to large datasets by using a two-step approach. First, MCMC is run in parallel on (sub)posteriors defined on data partitions. Then, a server combines local results. While efficient, this framework is very sensitive to the quality of subposterior sampling. Common sampling problems such as missing modes or misrepresentation of low-density regions are amplified – instead of being corrected – in the combination phase, leading to catastrophic failures. In this work, we propose a novel combination strategy to mitigate this issue. Our strategy, Parallel Active Inference (PAI), leverages Gaussian Process (GP) surrogate modeling and active learning. After fitting GPs to subposteriors, PAI (i) shares information between GP surrogates to cover missing modes; and (ii) uses active sampling to individually refine subposterior approximations. We validate PAI in challenging benchmarks, including heavy-tailed and multi-modal posteriors and a real-world application to computational neuroscience. Empirical results show that PAI succeeds where previous methods catastrophically fail, with a small communication overhead. Daniel Augusto R. M. A. de Souza, Diego Mesquita, Samuel Kaski, Luigi Acerbi |
AISTATS | 1 |
| 2022 | Self-tuning portfolio-based Bayesian optimization
Thiago de P. Vasconcelos, Daniel Augusto R. M. A. de Souza, Gustavo C. de M. Virgolino, César Lincoln C. Mattos, João Paulo Pordeus Gomes |
Expert Syst. Appl. | 2 |
| 2021 | Learning GPLVM with arbitrary kernels using the unscented transformationabstractGaussian Process Latent Variable Model (GPLVM) is a flexible framework to handle uncertain inputs in Gaussian Processes (GPs) and incorporate GPs as components of larger graphical models. Nonetheless, the standard GPLVM variational inference approach is tractable only for a narrow family of kernel functions. The most popular implementations of GPLVM circumvent this limitation using quadrature methods, which may become a computational bottleneck even for relatively low dimensions. For instance, the widely employed Gauss-Hermite quadrature has exponential complexity on the number of dimensions. In this work, we propose using the unscented transformation instead. Overall, this method presents comparable, if not better, performance than off-the-shelf solutions to GPLVM, and its computational complexity scales only linearly on dimension. In contrast to Monte Carlo methods, our approach is deterministic and works well with quasi-Newton methods, such as the Broyden-Fletcher-Goldfarb-Shanno (BFGS) algorithm. We illustrate the applicability of our method with experiments on dimensionality reduction and multistep-ahead prediction with uncertainty propagation. Daniel Augusto R. M. A. de Souza, Diego Mesquita, João Paulo Pordeus Gomes, César Lincoln C. Mattos |
AISTATS | 1 |
| 2019 | No-PASt-BO: Normalized Portfolio Allocation Strategy for Bayesian OptimizationabstractBayesian Optimization (BO) is a framework for black-box optimization that is especially suitable for expensive cost functions. Among the main parts of a BO algorithm, the acquisition function is of fundamental importance, since it guides the optimization algorithm by translating the uncertainty of the regression model in a utility measure for each point to be evaluated. Considering such aspect, selection and design of acquisition functions are one of the most popular research topics in BO. Since no single acquisition function was proved to have better performance in all tasks, a well-established approach consists of selecting different acquisition functions along the iterations of a BO execution. In such an approach, the GP-Hedge algorithm is a widely used option given its simplicity and good performance. Despite its success in various applications, GP-Hedge shows an undesirable characteristic of accounting on all past performance measures of each acquisition function to select the next function to be used. In this case, good or bad values obtained in an initial iteration may impact the choice of the acquisition function for the rest of the algorithm. This fact may induce a dominant behavior of an acquisition function and impact the final performance of the method. Aiming to overcome such limitation, in this work we propose a variant of GP-Hedge, named No-PASt-BO, that reduce the influence of far past evaluations. Moreover, our method presents a built-in normalization that avoids the functions in the portfolio to have similar probabilities, thus improving the exploration. The obtained results on both synthetic and real-world optimization tasks indicate that No-PASt-BO presents competitive performance and always outperforms GP-Hedge. Thiago de P. Vasconcelos, Daniel Augusto R. M. A. de Souza, César Lincoln C. Mattos, João Paulo Pordeus Gomes |
ICTAI | 2 |