EDBT 2026 Demo / reviewers in the wild / expert
Diego Mesquita
dblp:163/4293 · also Diego P. P. Mesquita, Diego Parente Paiva Mesquita
· DBLP profile ↗
31ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0002-9061-7041ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 10 first-author · 14 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Expert-aided causal discovery of ancestral graphsabstractPublisher Copyright: © 2026 The Author(s) | openaire: EC/H2020/951847/EU//ELISE Tiago da Silva, Bruna Bazaluk, Eliezer S. Silva, António Góis, Salem Lahlou, Dominik Heider, Samuel Kaski, Diego Mesquita, Adèle Helena Ribeiro |
Inf. Sci. | 8 |
| 2025 | When do GFlowNets learn the right distribution?abstractGenerative Flow Networks (GFlowNets) are an emerging class of sampling methods for distributions over discrete and compositional objects, e.g., graphs. In spite of their remarkable success in problems such as drug discovery and phylogenetic inference, the question of when and whether GFlowNets learn to sample from the target distribution remains underexplored. To tackle this issue, we first assess the extent to which a violation of the detailed balance of the underlying flow network might hamper the correctness of GFlowNet's sampling distribution. In particular, we demonstrate that the impact of an imbalanced edge on the model's accuracy is influenced by the total amount of flow passing through it and, as a consequence, is unevenly distributed across the network. We also argue that, depending on the parameterization, imbalance may be inevitable. In this regard, we consider the problem of sampling from distributions over graphs with GFlowNets parameterized by graph neural networks (GNNs) and show that the representation limits of GNNs delineate which distributions these GFlowNets can approximate. Lastly, we address these limitations by proposing a theoretically sound and computationally tractable metric for assessing GFlowNets, experimentally showing it is a better proxy for correctness than popular evaluation protocols. Tiago da Silva, Rodrigo Barreto Alves, Eliezer S. Silva, Amauri H. Souza, Vikas Garg 0001, Samuel Kaski, Diego Mesquita |
ICLR | 7 |
| 2025 | Generalization and Distributed Learning of GFlowNetsabstractConventional wisdom attributes the success of Generative Flow Networks (GFlowNets) to their ability to exploit the compositional structure of the sample space for learning generalizable flow functions (Bengio et al., 2021). Despite the abundance of empirical evidence, formalizing this belief with verifiable non-vacuous statistical guarantees has remained elusive. We address this issue with the first data-dependent generalization bounds for GFlowNets. We also elucidate the negative impact of the state space size on the generalization performance of these models via Azuma-Hoeffding-type oracle PAC-Bayesian inequalities. We leverage our theoretical insights to design a novel distributed learning algorithm for GFlowNets, which we call *Subgraph Asynchronous Learning* (SAL). In a nutshell, SAL utilizes a divide-and-conquer strategy: multiple GFlowNets are trained in parallel on smaller subnetworks of the flow network, and then aggregated with an additional GFlowNet that allocates appropriate flow to each subnetwork. Our experiments with synthetic and real-world problems demonstrate the benefits of SAL over centralized training in terms of mode coverage and distribution matching. Amauri H. Souza, Omar Rivasplata, Vikas Garg 0001, Samuel Kaski, Diego Mesquita |
ICLR | 6 |
| 2025 | Infinite Neural Operators: Gaussian processes on functionsabstractA variety of infinitely wide neural architectures (e.g., dense NNs, CNNs, and transformers) induce Gaussian process (GP) priors over their outputs.
These relationships provide both an accurate characterization of the prior predictive distribution and enable the use of GP machinery to improve the uncertainty quantification of deep neural networks.
In this work, we extend this connection to neural operators (NOs), a class of models designed to learn mappings between function spaces.
Specifically, we show conditions for when arbitrary-depth NOs with Gaussian-distributed convolution kernels converge to function-valued GPs.
Based on this result, we show how to compute the covariance functions of these NO-GPs for two NO parametrizations, including the popular Fourier neural operator (FNO).
With this, we compute the posteriors of these GPs in regression scenarios, including PDE solution operators.
This work is an important step towards uncovering the inductive biases of current FNO architectures and opens a path to incorporate novel inductive biases for use in kernel-based operator learning methods. Daniel Augusto R. M. A. de Souza, Jake Cunningham, Yuri F. Saporito, Diego Mesquita, Marc Peter Deisenroth |
NeurIPS | 5 |
| 2025 | Differentially Private Selection Using Smooth SensitivityabstractDifferentially private selection mechanisms offer strong privacy guarantees for queries aiming to identify the top-scoring element$r$from a finite set$\mathfrak{R}$, based on a dataset-dependent utility function. While selection queries are fundamental in data science, few mechanisms effectively ensure their privacy. Furthermore, most approaches rely on global sensitivity to achieve differential privacy (DP), which can introduce excessive noise and impair downstream inferences. To address this limitation, we propose the Smooth Noisy Max (SNM) mechanism, which leverages smooth sensitivity to yield provably tighter (upper bounds on) expected errors compared to global sensitivity-based methods. Empirical results demonstrate that SNM is more accurate than state-of-the-art differentially private selection methods in three applications: percentile selection, greedy decision trees, and random forests. Iago C. Chaves, Victor A. E. de Farias, Amanda Perez, Diego Mesquita, Javam C. Machado |
SP | 4 |
| 2024 | Amortized Variational Deep Kernel LearningabstractDeep kernel learning (DKL) marries the uncertainty quantification of Gaussian processes (GPs) and the representational power of deep neural networks. However, training DKL is challenging and often leads to overfitting. Most notably, DKL often learns “non-local” kernels — incurring spurious correlations. To remedy this issue, we propose using amortized inducing points and a parameter-sharing scheme, which ties together the amortization and DKL networks. This design imposes an explicit dependency between the ELBO’s model fit and capacity terms. In turn, this prevents the former from dominating the optimization procedure and incurring the aforementioned spurious correlations. Extensive experiments show that our resulting method, amortized varitional DKL (AVDKL), i) consistently outperforms DKL and standard GPs for tabular data; ii) achieves significantly higher accuracy than DKL in node classification tasks; and iii) leads to substantially better accuracy and negative log-likelihood than DKL on CIFAR100. Alan Lucas Silva Matias, César Lincoln C. Mattos, João Paulo Pordeus Gomes, Diego Mesquita |
ICML | 4 |
| 2024 | Embarrassingly Parallel GFlowNetsabstractGFlowNets are a promising alternative to MCMC sampling for discrete compositional random variables. Training GFlowNets requires repeated evaluations of the unnormalized target distribution, or reward function. However, for large-scale posterior sampling, this may be prohibitive since it incurs traversing the data several times. Moreover, if the data are distributed across clients, employing standard GFlowNets leads to intensive client-server communication. To alleviate both these issues, we propose embarrassingly parallel GFlowNet (EP-GFlowNet). EP-GFlowNet is a provably correct divide-and-conquer method to sample from product distributions of the form $R(\cdot) \propto R_1(\cdot) ... R_N(\cdot)$ — e.g., in parallel or federated Bayes, where each $R_n$ is a local posterior defined on a data partition. First, in parallel, we train a local GFlowNet targeting each $R_n$ and send the resulting models to the server. Then, the server learns a global GFlowNet by enforcing our newly proposed aggregating balance condition, requiring a single communication step. Importantly, EP-GFlowNets can also be applied to multi-objective optimization and model reuse. Our experiments illustrate the effectiveness of EP-GFlowNets on multiple tasks, including parallel Bayesian phylogenetics, multi-objective multiset and sequence generation, and federated Bayesian structure learning. Tiago da Silva, Luiz Max Carvalho, Amauri H. Souza, Samuel Kaski, Diego Mesquita |
ICML | 5 |
| 2024 | Streaming Bayes GFlowNetsabstractBayes' rule naturally allows for inference refinement in a streaming fashion, without the need to recompute posteriors from scratch whenever new data arrives. In principle, Bayesian streaming is straightforward: we update our prior with the available data and use the resulting posterior as a prior when processing the next data chunk. In practice, however, this recipe entails i) approximating an intractable posterior at each time step; and ii) encapsulating results appropriately to allow for posterior propagation. For continuous state spaces, variational inference (VI) is particularly convenient due to its scalability and the tractability of variational posteriors, For discrete state spaces, however, state-of-the-art VI results in analytically intractable approximations that are ill-suited for streaming settings. To enable streaming Bayesian inference over discrete parameter spaces, we propose streaming Bayes GFlowNets (abbreviated as SB-GFlowNets) by leveraging the recently proposed GFlowNets --- a powerful class of amortized samplers for discrete compositional objects. Notably, SB-GFlowNet approximates the initial posterior using a standard GFlowNet and subsequently updates it using a tailored procedure that requires only the newly observed data. Our case studies in linear preference learning and phylogenetic inference showcase the effectiveness of SB-GFlowNets in sampling from an unnormalized posterior in a streaming setting. As expected, we also observe that SB-GFlowNets is significantly faster than repeatedly training a GFlowNet from scratch to sample from the full posterior. Tiago da Silva, Daniel Augusto R. M. A. de Souza, Diego Mesquita |
NeurIPS | 3 |
| 2024 | On Divergence Measures for Training GFlowNetsabstractGenerative Flow Networks (GFlowNets) are amortized samplers of unnormalized distributions over compositional objects with applications to causal discovery, NLP, and drug design. Recently, it was shown that GFlowNets can be framed as a hierarchical variational inference (HVI) method for discrete distributions. Despite this equivalence, attempts to train GFlowNets using traditional divergence measures as learning objectives were unsuccessful. Instead, current approaches for training these models rely on minimizing the log-squared difference between a proposal (forward policy) and a target (backward policy) distributions. In this work, we first formally extend the relationship between GFlowNets and HVI to distributions on arbitrary measurable topological spaces. Then, we empirically show that the ineffectiveness of divergence-based learning of GFlowNets is due to large gradient variance of the corresponding stochastic objectives. To address this issue, we devise a collection of provably variance-reducing control variates for gradient estimation based on the REINFORCE leave-one-out estimator. Our experimental results suggest that the resulting algorithms often accelerate training convergence when compared against previous approaches. All in all, our work contributes by narrowing the gap between GFlowNet training and HVI, paving the way for algorithmic advancements inspired by the divergence minimization viewpoint. Tiago da Silva, Eliezer S. Silva, Diego Mesquita |
NeurIPS | 3 |
| 2024 | Towards automatic labeling of exception handling bugs: A case study of 10 years bug-fixing in Apache Hadoop
Antônio da Silva, Renan Gomes Vieira, Diego Mesquita, João Paulo Pordeus Gomes, Lincoln S. Rocha |
Empir. Softw. Eng. | 3 |
| 2023 | Distill n' Explain: explaining graph neural networks using simple surrogatesabstractExplaining node predictions in graph neural networks (GNNs) often boils down to finding graph substructures that preserve predictions. Finding these structures usually implies back-propagating through the GNN, bonding the complexity (e.g., number of layers) of the GNN to the cost of explaining it. This naturally begs the question: Can we break this bond by explaining a simpler surrogate GNN? To answer the question, we propose Distill n’ Explain (DnX). First, DnX learns a surrogate GNN via knowledge distillation. Then, DnX extracts node or edge-level explanations by solving a simple convex program. We also propose FastDnX, a faster version of DnX that leverages the linear decomposition of our surrogate model. Experiments show that DnX and FastDnX often outperform state-of-the-art GNN explainers while being orders of magnitude faster. Additionally, we support our empirical findings with theoretical results linking the quality of the surrogate model (i.e., distillation error) to the faithfulness of explanations. Tamara A. Pereira, Erik Jhones F. do Nascimento, Lucas Resck, Diego Mesquita, Amauri H. Souza |
AISTATS | 4 |
| 2023 | Thin and deep Gaussian processesabstractGaussian processes (GPs) can provide a principled approach to uncertainty quantification with easy-to-interpret kernel hyperparameters, such as the lengthscale, which controls the correlation distance of function values.However, selecting an appropriate kernel can be challenging.
Deep GPs avoid manual kernel engineering by successively parameterizing kernels with GP layers, allowing them to learn low-dimensional embeddings of the inputs that explain the output data.
Following the architecture of deep neural networks, the most common deep GPs warp the input space layer-by-layer but lose all the interpretability of shallow GPs. An alternative construction is to successively parameterize the lengthscale of a kernel, improving the interpretability but ultimately giving away the notion of learning lower-dimensional embeddings. Unfortunately, both methods are susceptible to particular pathologies which may hinder fitting and limit their interpretability.
This work proposes a novel synthesis of both previous approaches: {Thin and Deep GP} (TDGP). Each TDGP layer defines locally linear transformations of the original input data maintaining the concept of latent embeddings while also retaining the interpretation of lengthscales of a kernel. Moreover, unlike the prior solutions, TDGP induces non-pathological manifolds that admit learning lower-dimensional representations.
We show with theoretical and experimental results that i) TDGP is, unlike previous models, tailored to specifically discover lower-dimensional manifolds in the input data, ii) TDGP behaves well when increasing the number of layers, and iii) TDGP performs well in standard benchmark datasets. Daniel Augusto R. M. A. de Souza, Alexander Nikitin 0002, St John, Magnus Ross, Mauricio A. Álvarez, Marc Peter Deisenroth, João Paulo Pordeus Gomes, Diego Mesquita, César Lincoln C. Mattos |
NeurIPS | 8 |
| 2022 | Parallel MCMC Without Embarrassing FailuresabstractEmbarrassingly parallel Markov Chain Monte Carlo (MCMC) exploits parallel computing to scale Bayesian inference to large datasets by using a two-step approach. First, MCMC is run in parallel on (sub)posteriors defined on data partitions. Then, a server combines local results. While efficient, this framework is very sensitive to the quality of subposterior sampling. Common sampling problems such as missing modes or misrepresentation of low-density regions are amplified – instead of being corrected – in the combination phase, leading to catastrophic failures. In this work, we propose a novel combination strategy to mitigate this issue. Our strategy, Parallel Active Inference (PAI), leverages Gaussian Process (GP) surrogate modeling and active learning. After fitting GPs to subposteriors, PAI (i) shares information between GP surrogates to cover missing modes; and (ii) uses active sampling to individually refine subposterior approximations. We validate PAI in challenging benchmarks, including heavy-tailed and multi-modal posteriors and a real-world application to computational neuroscience. Empirical results show that PAI succeeds where previous methods catastrophically fail, with a small communication overhead. Daniel Augusto R. M. A. de Souza, Diego Mesquita, Samuel Kaski, Luigi Acerbi |
AISTATS | 2 |
| 2022 | Bayesian Analysis of Bug-Fixing Time using Report DataabstractBackground: Bug-fixing is the crux of software maintenance. It entails tending to heaps of bug reports using limited resources. Using historical data, we can ask questions that contribute to better-informed allocation heuristics. The caveat here is that often there is not enough data to provide a sound response. This issue is especially prominent for young projects. Also, answers may vary from project to project. Consequently, it is impossible to generalize results without assuming a notion of relatedness between projects. Renan Gomes Vieira, Diego Mesquita, César Lincoln C. Mattos, Ricardo Britto 0001, Lincoln S. Rocha, João Paulo Pordeus Gomes |
ESEM | 2 |
| 2022 | Provably expressive temporal graph networksabstractTemporal graph networks (TGNs) have gained prominence as models for embedding dynamic interactions, but little is known about their theoretical underpinnings. We establish fundamental results about the representational power and limits of the two main categories of TGNs: those that aggregate temporal walks (WA-TGNs), and those that augment local message passing with recurrent memory modules (MP-TGNs). Specifically, novel constructions reveal the inadequacy of MP-TGNs and WA-TGNs, proving that neither category subsumes the other. We extend the 1-WL (Weisfeiler-Leman) test to temporal graphs, and show that the most powerful MP-TGNs should use injective updates, as in this case they become as expressive as the temporal WL. Also, we show that sufficiently deep MP-TGNs cannot benefit from memory, and MP/WA-TGNs fail to compute graph properties such as girth. These theoretical insights lead us to PINT --- a novel architecture that leverages injective temporal message passing and relative positional features. Importantly, PINT is provably more expressive than both MP-TGNs and WA-TGNs. PINT significantly outperforms existing TGNs on several real-world benchmarks. Amauri H. Souza, Diego Mesquita, Samuel Kaski, Vikas Garg 0001 |
NeurIPS | 2 |
| 2022 | Bayesian MultilaterationabstractMultilateration (MLAT) is thede factotechnique to localize points of interest (POIs) in navigation and surveillance systems. Despite sensors being inherently noisy, most existing techniques i) are oblivious to noise patterns in sensor measurements; and ii) only provide point estimates of the POI. This often results in unreliable estimates with high variance,i.e., that are highly sensitive to measurement noise. To overcome this caveat, we advocate the use of Bayesian modeling. Using Bayesian statistics, we provide a comprehensive guide to handle uncertainties in MLAT, including principled choices for the likelihood function and the prior distributions. Notably, the resulting model is easy to implement and can leverage off-the-shelf Markov Chain Monte Carlo (MCMC) software for inference. Besides coping with unreliable measurements, our framework can also deal with sensors whose location is not completely known, which is an asset in mobile systems. Our solution also naturally incorporates multiple measurements per reference point, a common practical situation that is usually not handled directly by other approaches. Comprehensive experiments with both synthetic and real-world data indicate that our Bayesian approach to the MLAT task provides better position estimation and uncertainty quantification when compared to the available alternatives. Alisson S. C. Alencar, César Lincoln C. Mattos, João Paulo Pordeus Gomes, Diego Mesquita |
IEEE Signal Process. Lett. | 4 |
| 2021 | Learning GPLVM with arbitrary kernels using the unscented transformationabstractGaussian Process Latent Variable Model (GPLVM) is a flexible framework to handle uncertain inputs in Gaussian Processes (GPs) and incorporate GPs as components of larger graphical models. Nonetheless, the standard GPLVM variational inference approach is tractable only for a narrow family of kernel functions. The most popular implementations of GPLVM circumvent this limitation using quadrature methods, which may become a computational bottleneck even for relatively low dimensions. For instance, the widely employed Gauss-Hermite quadrature has exponential complexity on the number of dimensions. In this work, we propose using the unscented transformation instead. Overall, this method presents comparable, if not better, performance than off-the-shelf solutions to GPLVM, and its computational complexity scales only linearly on dimension. In contrast to Monte Carlo methods, our approach is deterministic and works well with quasi-Newton methods, such as the Broyden-Fletcher-Goldfarb-Shanno (BFGS) algorithm. We illustrate the applicability of our method with experiments on dimensionality reduction and multistep-ahead prediction with uncertainty propagation. Daniel Augusto R. M. A. de Souza, Diego Mesquita, João Paulo Pordeus Gomes, César Lincoln C. Mattos |
AISTATS | 2 |
| 2021 | Improving Graph Variational Autoencoders with Multi-Hop Simple ConvolutionsabstractVariational auto-encoding architectures represent one of the most popular approaches to graph generative modeling.These models comprise encoder and a decoder networks, which map back and forth between the input and latent spaces.Notably, most of the literature in variational autoencoders (VAEs) for graphs focuses on developing more efficient architectures at the expense of increased complexity.In this work, we pursue an orthogonal direction and leverage multi-hop linear graph convolutional layers to create efficient yet simple encoders, boosting the performance of graph autoencoders.Our results demonstrate that our approach outperforms popular graph VAE baselines in link prediction tasks. Erik Jhones F. do Nascimento, Amauri H. Souza, Diego Mesquita |
ESANN | 3 |
| 2021 | Federated stochastic gradient Langevin dynamicsabstractStochastic gradient MCMC methods, such as stochastic gradient Langevin dynamics (SGLD), employ fast but noisy gradient estimates to enable large-scale posterior sampling. Although we can easily extend SGLD to distributed settings, it suffers from two issues when applied to federated non-IID data. First, the variance of these estimates increases significantly. Second, delaying communication causes the Markov chains to diverge from the true posterior even for very simple models. To alleviate both these problems, we propose conducive gradients, a simple mechanism that combines local likelihood approximations to correct gradient updates. Notably, conducive gradients are easy to compute, and since we only calculate the approximations once, they incur negligible overhead. We apply conducive gradients to distributed stochastic gradient Langevin dynamics (DSGLD) and call the resulting method “federated stochastic gradient Langevin dynamics” (FSGLD). We demonstrate that our approach can handle delayed communication rounds, converging to the target posterior in cases where DSGLD fails. We also show that FSGLD outperforms DSGLD for non-IID federated data with experiments on metric learning and neural networks. Khaoula el Mekkaoui, Diego Mesquita, Paul Blomstedt, Samuel Kaski |
UAI | 2 |
| 2020 | Rethinking pooling in graph neural networksabstractGraph pooling is a central component of a myriad of graph neural network (GNN) architectures. As an inheritance from traditional CNNs, most approaches formulate graph pooling as a cluster assignment problem, extending the idea of local patches in regular grids to graphs. Despite the wide adherence to this design choice, no work has rigorously evaluated its influence on the success of GNNs. In this paper, we build upon representative GNNs and introduce variants that challenge the need for locality-preserving representations, either using randomization or clustering on the complement graph. Strikingly, our experiments demonstrate that using these variants does not result in any decrease in performance. To understand this phenomenon, we study the interplay between convolutional layers and the subsequent pooling ones. We show that the convolutions play a leading role in the learned representations. In contrast to the common belief, local pooling is not responsible for the success of GNNs on relevant and widely-used benchmarks. Diego Mesquita, Amauri H. Souza, Samuel Kaski |
NeurIPS | 1 |
| 2020 | A sparse linear regression model for incomplete datasets
Marcelo B. A. Veras, Diego Mesquita, César Lincoln C. Mattos, João Paulo Pordeus Gomes |
Pattern Anal. Appl. | 2 |
| 2020 | LS-SVR as a Bayesian RBF Network
Diego Mesquita, Luis A. Freitas, João Paulo Pordeus Gomes, César Lincoln C. Mattos |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Embarrassingly Parallel MCMC using Deep Invertible Transformations
Diego Mesquita, Paul Blomstedt, Samuel Kaski |
UAI | 1 |
| 2019 | Artificial Neural Networks with Random Weights for Incomplete Datasets
Diego Mesquita, João Paulo Pordeus Gomes, Leonardo Ramos Rodrigues |
Neural Process. Lett. | 1 |
| 2017 | A Robust Minimal Learning Machine based on the M-Estimator
João Paulo Pordeus Gomes, Diego Mesquita, Ananda Freire, Amauri H. Souza, Tommi Kärkkäinen |
ESANN | 2 |
| 2017 | Euclidean distance estimation in incomplete datasets
Diego Mesquita, João Paulo Pordeus Gomes, Amauri H. Souza, Juvêncio S. Nobre |
Neurocomputing | 1 |
| 2017 | Ensemble of Efficient Minimal Learning Machines for Classification and Regression
Diego Mesquita, João Paulo Pordeus Gomes, Amauri H. Souza |
Neural Process. Lett. | 1 |
| 2016 | K-means for Datasets with Missing Attributes: Building Soft Constraints with Observed and Imputed Values
Diego Mesquita, João Paulo Pordeus Gomes, Leonardo Ramos Rodrigues |
ESANN | 1 |
| 2016 | Using Robust Extreme Learning Machines to Predict Cotton Yarn Strength and Hairiness
Diego Mesquita, Antônio C. Araújo Neto, Jose Queiroz Neto, João Paulo Pordeus Gomes, Leonardo Ramos Rodrigues |
ESANN | 1 |
| 2016 | Radial Basis Function Neural Networks for Datasets with Missing Values
Diego Mesquita, João Paulo Pordeus Gomes |
ISDA | 1 |
| 2015 | A Minimal Learning Machine for Datasets with Missing Values
Diego Mesquita, João Paulo Pordeus Gomes, Amauri H. Souza |
ICONIP (1) | 1 |