Victor Elvira

dblp:63/7883 · also Víctor Elvira · DBLP profile ↗
← Back
59ranked-venue papers
9as first author
33since 2021 · last 2026
0000-0002-8967-4866ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 9 first-author · 16 since 2021Artificial intelligence and machine learning · 12 · 12 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Differentiable interacting multiple model particle filtering
John-Joseph Brady, Yuhui Luo, Wenwu Wang 0001, Victor Elvira, Yunpeng Li 0001
Signal Process.4
2025 Learning a Sparse Polynomial Approximation to the Transition Function of General State-Space Models
abstract
State-space models are a statistical framework for modelling temporal phenomena via a hidden state. In this framework, the hidden state is not observed, and instead a series of related observations are obtained. A state-space model is defined by the state dynamics, which is encoded as a distribution. The parameters of this distribution are often unknown, and must be estimated in order to perform inference. In real-world systems, it is common that not all dimensions of the hidden state directly interact, which implies a sparse system. Most parameter estimation methods for state-space models cannot recover sparsity. In this work, we propose PolyGrad, a fully automatic method for obtaining sparse estimates of the state interactions of a non-linear state-space model via a polynomial approximation. This novel versatile methodology allows to infer both the structure and values of a generic parameterisation of a state-space model. The proposed method is computationally efficient and can represent a large class of complex systems.
Benjamin Cox, Emilie Chouzenoux, Victor Elvira
ICASSP3
2025 Subset Adaptive Importance Sampling for Multi-Failure-Region Estimation
abstract
Rare-event estimation is a relevant problem in statistical signal processing. While significant efforts have been devoted to unimodal problems, the case of multimodal scenarios is still hard to tackle with state-of-the-art methods. In this paper, we propose the subset adaptive importance sampling (SAIS) algorithm for the estimation of rare events in the context of Bayesian signal processing. The new algorithm incorporates advantages from subset simulation techniques and from adaptive importance sampling, in particular population Monte Carlo methods. It provides a multilevel estimation of the failure probability and balances the trade-off between exploration and exploitation of the failure space in the adaptation. The algorithm is particularly well suited for multimodal scenarios. We further introduce adaptation rules for more efficient sampling. Finally, we test the good performance of the proposed algorithm in numerical examples.
Sara Helal, Victor Elvira
ICASSP2
2025 Dimensionality Reduction with Entropies from f-Divergences
Mateu Sbert, Min Chen 0001, Jordi Poch, Miquel Feixas, Shuning Chen, Victor Elvira
MDAI6
2025 Regularized Rényi Divergence Minimization through Bregman Proximal Gradient Algorithms
abstract
We study the variational inference problem of minimizing a regularized Rényi divergence over an exponential family. We propose to solve this problem with a Bregman proximal gradient algorithm. We propose a sampling-based algorithm to cover the black-box setting, corresponding to a stochastic Bregman proximal gradient algorithm with biased gradient estimator. We show that the resulting algorithms can be seen as relaxed moment-matching algorithms with an additional proximal step. Using Bregman updates instead of Euclidean ones allows us to exploit the geometry of our approximate model. We prove strong convergence guarantees for both our deterministic and stochastic algorithms using this viewpoint, including monotonic decrease of the objective, convergence to a stationary point or to the minimizer, and geometric convergence rates. These new theoretical insights lead to a versatile, robust, and competitive method, as illustrated by numerical experiments
Thomas Guilmeau, Emilie Chouzenoux, Victor Elvira
J. Mach. Learn. Res.3
2025 Learning state and proposal dynamics in state-space models using differentiable particle filters and neural networks
abstract
State-space models are a popular statistical framework for analysing sequential data. Within this framework, particle filters are often used to perform inference on non-linear state-space models. We introduce a new method, StateMixNN, that uses a pair of neural networks to learn the proposal distribution and transition kernel of a particle filter. Both distributions are approximated using multivariate Gaussian mixtures. The component means and covariances of these mixtures are learnt as outputs of learned functions. Our method is trained targeting the log-likelihood, thereby requiring only the observation series, and combines the interpretability of state-space models with the flexibility and approximation power of artificial neural networks . The proposed method significantly improves recovery of the hidden state in comparison with the state-of-the-art, showing greater improvement in highly non-linear scenarios.
Benjamin Cox, Santiago Segarra, Victor Elvira
Signal Process.3
2025 A Proximal Newton Adaptive Importance Sampler
abstract
Adaptive importance sampling (AIS) algorithms are a rising methodology in signal processing, statistics, and machine learning. An effective adaptation of the proposals is key for the success of AIS. Recent works have shown that gradient information about the involved target density can greatly boost performance, but its applicability is restricted to differentiable targets. In this letter, we propose a proximal Newton adaptive importance sampler for the estimation of expectations with respect to non-smooth target distributions. We implement a scaled Newton proximal gradient method to adapt the proposal distributions, enabling efficient and optimized moves even when the target distribution lacks differentiability. We show the good performance of the algorithm in two scenarios: one with convex constraints and another with non-smooth sparse priors.
Victor Elvira, Emilie Chouzenoux, Ömer Deniz Akyildiz
IEEE Signal Process. Lett.1
2025 A Divergence-Based Condition to Ensure Quantile Improvement in Black-Box Global Optimization
abstract
black-box global optimization aims at minimizing an objective function whose analytical form is not known. To do so, many state-of-the-art methods rely on sampling-based strategies, where sampling distributions are built in an iterative fashion, so that their mass concentrate where the objective function is low. Despite empirical success, the theoretical study of these methods remains difficult. In this work, we introduce a new framework, based on divergence-decrease conditions, to study and design black-box global optimization algorithms. Our approach allows to establish and quantify the improvement of sampling distributions at each iteration, in terms of expected value or quantile of the objective. We show that the information-geometric optimization approach fits within our framework, yielding a new approach for its analysis. We also establish sampling distribution improvement results for two novel algorithms, one related with the cross-entropy approach with mixture models, and another one using heavy-tailed sampling distributions.
Thomas Guilmeau, Emilie Chouzenoux, Victor Elvira
IEEE Trans. Evol. Comput.3
2024 Adaptive importance sampling for heavy-tailed distributions via α-divergence minimization
abstract
Adaptive importance sampling (AIS) algorithms are widely used to approximate expectations with respect to complicated target probability distributions. When the target has heavy tails, existing AIS algorithms can provide inconsistent estimators or exhibit slow convergence, as they often neglect the target’s tail behaviour. To avoid this pitfall, we propose an AIS algorithm that approximates the target by Student-t proposal distributions. We adapt location and scale parameters by matching the escort moments - which are defined even for heavy-tailed distributions - of the target and proposal. These updates minimize the $\alpha$-divergence between the target and the proposal, thereby connecting with variational inference. We then show that the $\alpha$-divergence can be approximated by a generalized notion of effective sample size and leverage this new perspective to adapt the tail parameter with Bayesian optimization. We demonstrate the efficacy of our approach through applications to synthetic targets and a Bayesian Student-t regression task on a real example with clinical trial data.
Thomas Guilmeau, Nicola Branchini, Emilie Chouzenoux, Victor Elvira
AISTATS4
2024 Variational Resampling
abstract
We cast the resampling step in particle filters (PFs) as a variational inference problem, resulting in a new class of resampling schemes: variational resampling. Variational resampling is flexible as it allows for choices of 1) divergence to minimize, 2) target distribution to input to the divergence, and 3) divergence minimization algorithm. With this novel application of VI to particle filters, variational resampling further unifies these two powerful and popular methodologies. We construct two variational resamplers that replicate particles in order to maximize lower bounds with respect to two different target measures. We benchmark our variational resamplers on challenging smoothing tasks, outperforming PFs that implement the state-of-the-art resampling schemes.
Oskar Kviman, Nicola Branchini, Victor Elvira, Jens Lagergren
AISTATS3
2024 Regime Learning for Differentiable Particle Filters
abstract
Differentiable particle filters are an emerging class of models that combine sequential Monte Carlo techniques with the flexibility of neural networks to perform state space inference. This paper concerns the case where the system may switch between a finite set of state-space models, i.e. regimes. No prior approaches effectively learn both the individual regimes and the switching process simultaneously. In this paper, we propose the neural network based regime learning differentiable particle filter (RLPF) to address this problem. We further design a training procedure for the RLPF and other related algorithms. We demonstrate competitive performance compared to the previous state-of-the-art algorithms on a pair of numerical experiments.
John-Joseph Brady, Yuhui Luo, Wenwu Wang 0001, Victor Elvira, Yunpeng Li 0001
FUSION4
2024 Graphical Inference in Non-Markovian Linear-Gaussian State-Space Models
abstract
State-space models (SSMs) are common tools in time-series analysis for inference and prediction. SSMs are versatile probabilistic models that allow for Bayesian inference by describing a (generally Markovian) latent process. However, the parameters of that latent process are often unknown and must be estimated. In this paper, we consider the parameter estimation in a SSM with a non-Markovian linear-Gaussian latent process. This process is described as a vector auto-regressive with p unknown matrices. Our algorithm LaGrangEM estimates these matrices through an expectation-maximization algorithm that exploits a graphical interpretation of the latent process in order to define prior knowledge about the unknown parameters. We connect the new algorithm with existing approaches such as Granger causality and graphical inference in SSMs. We discuss the strong potential of the algorithm to bring interpretability, e.g., in estimating causal relationships and their delays. The numerical experiments also show a superiority in performance.
Emilie Chouzenoux, Victor Elvira
ICASSP2
2024 End-to-End Learning of Gaussian Mixture Proposals Using Differentiable Particle Filters and Neural Networks
abstract
We introduce a new method, named PropMixNN, that uses a neural network to learn the proposal distribution of a particle filter. The optimal proposal distribution is approximated as a multivariate Gaussian mixture, so the proposed method aims at learning the means and covariance matrices of the S components that characterise the mixture. This unsupervised method is trained to target the log-likelihood, which does not require knowledge of the hidden state. The performance of the method is assessed in a stochastic Lorenz 96 model, which presents a non-linear chaotic behaviour. The proposed method reduces estimation errors in comparison with the state-of-the-art, showing greater improvement in highly non-linear scenarios.
Benjamin Cox, Sara Pérez-Vieites, Nicolas Zilberstein, Martin Sevilla, Santiago Segarra, Victor Elvira
ICASSP6
2024 Efficient Mixture Learning in Black-Box Variational Inference
abstract
Mixture variational distributions in black box variational inference (BBVI) have demonstrated impressive results in challenging density estimation tasks. However, currently scaling the number of mixture components can lead to a linear increase in the number of learnable parameters and a quadratic increase in inference time due to the evaluation of the evidence lower bound (ELBO). Our two key contributions address these limitations. First, we introduce the novel Multiple Importance Sampling Variational Autoencoder (MISVAE), which amortizes the mapping from input to mixture-parameter space using one-hot encodings. Fortunately, with MISVAE, each additional mixture component incurs a negligible increase in network parameters. Second, we construct two new estimators of the ELBO for mixtures in BBVI, enabling a tremendous reduction in inference time with marginal or even improved impact on performance. Collectively, our contributions enable scalability to hundreds of mixture components and provide superior estimation performance in shorter time, with fewer network parameters compared to previous Mixture VAEs. Experimenting with MISVAE, we achieve astonishing, SOTA results on MNIST. Furthermore, we empirically validate our estimators in other BBVI settings, including Bayesian phylogenetic inference, where we improve inference times for the SOTA mixture model on eight data sets.
Alexandra Hotti, Oskar Kviman, Ricky Molén, Victor Elvira, Jens Lagergren
ICML4
2024 WALGREEN: Web Based Platform for Soil Organic Carbon Inference Applications
abstract
Remote sensing data management and its use for classification and inference purposes is at the forefront of research tasks nowadays. There are, however, some inherent drawbacks and difficulties when dealing with, and understanding how satellite information is provided (particularly when referring to multiband/multispectral satellite platforms) and how different and disparate datasets related to soil content can be used and merged with this imagery.We present WALGREEN. The aim of this tool is to provide a secure environment to handle the whole process to use polygons or, geographical coordinates in tiff/geotiff images, have real-time access to images, save and get soil organic carbon real measurements, and generate datasets for machine learning training and inferential methods. We also aim to providing a framework to preprocess soil organic carbon information from different but accepted sources, like the Land Use/Cover Area frame statistical Survey database, so that even without real measurements, researchers may be able to start training different machine learning methodologies.
José Manuel Aroca, José-Francisco Díez-Pastor, Pedro Latorre-Carmona, Antonio Canepa-Oneto, Juan Carlos Rad, Gustau Camps-Valls, Victor Elvira, César Ignacio García-Osorio
IGARSS7
2024 Sparse Graphical Linear Dynamical Systems
abstract
Time-series datasets are central in machine learning with applications in numerous fields of science and engineering, such as biomedicine, Earth observation, and network analysis. Extensive research exists on state-space models (SSMs), which are powerful mathematical tools that allow for probabilistic and interpretable learning on time series. Learning the model parameters in SSMs is arguably one of the most complicated tasks, and the inclusion of prior knowledge is known to both ease the interpretation but also to complicate the inferential tasks. Very recent works have attempted to incorporate a graphical perspective on some of those model parameters, but they present notable limitations that this work addresses. More generally, existing graphical modeling tools are designed to incorporate either static information, focusing on statistical dependencies among independent random variables (e.g., graphical Lasso approach), or dynamic information, emphasizing causal relationships among time series samples (e.g., graphical Granger approaches). However, there are no joint approaches combining static and dynamic graphical modeling within the context of SSMs. This work proposes a novel approach to fill this gap by introducing a joint graphical modeling framework that bridges the graphical Lasso model and a causal-based graphical approach for the linear-Gaussian SSM. We present DGLASSO (Dynamic Graphical Lasso), a new inference method within this framework that implements an efficient block alternating majorization-minimization algorithm. The algorithm's convergence is established by departing from modern tools from nonlinear analysis. Experimental validation on various synthetic data showcases the effectiveness of the proposed model and inference algorithm. This work will significantly contribute to the understanding and utilization of time-series data in diverse scientific and engineering applications where incorporating a graphical approach is essential to perform the inference.
Emilie Chouzenoux, Victor Elvira
J. Mach. Learn. Res.2
2024 Hierarchical Average Fusion With GM-PHD Filters Against FDI and DoS Attacks
abstract
We address the multisensor multitarget tracking problem based on a hierarchical sensor network. In this setup, there is a fusion center, several cluster heads, and many sensors. Each sensor runs a Gaussian mixture probability hypothesis density (PHD) filter. The sensors send their locally calculated Gaussian components to the local cluster head in the presence of false data injection (FDI) and denial-of-service (DoS) attackers. We propose a hybrid PHD averaging fusion framework that consists of two parts: one uses the arithmetic average (AA) fusion to compensate for information shortage due to DoS and the other uses the geometric average (GA) fusion to suppress false information due to FDI. By integrating the respective zero forcing and avoiding behaviors of the two average fusion approaches, our proposed hybrid fusion scheme is proven resilient to both FDI and DoS attacks. Experimental results illustrate that our proposed algorithm can provide reliable tracking performance against FDI and DoS attacks.
Tiancheng Li 0002, Junkun Yan, Victor Elvira
IEEE Signal Process. Lett.4
2023 Graphit: Iterative Reweighted ℓ1 Algorithm for Sparse Graph Inference in State-Space Models
abstract
State-space models (SSMs) are a common tool for modeling multi-variate discrete-time signals. The linear-Gaussian (LG) SSM is widely applied as it allows for a closed-form solution at inference, if the model parameters are known. However, they are rarely available in real-world problems and must be estimated. Promoting sparsity of these parameters favours both interpretability and tractable inference. In this work, we propose GraphIT, a majorization-minimization (MM) algorithm for estimating the linear operator in the state equation of an LG-SSM under sparse prior. A versatile family of non-convex regularization potentials is proposed. The MM method relies on tools inherited from the expectation-maximization methodology and the iterated reweighted-l1 approach. In particular, we derive a suitable convex upper bound for the objective function, that we then minimize using a proximal splitting algorithm. Numerical experiments illustrate the benefits of the proposed inference technique.
Emilie Chouzenoux, Victor Elvira
ICASSP2
2023 Adaptive Simulated Annealing Through Alternating Rényi Divergence Minimization
abstract
Simulated annealing is a popular approach to solve nonconvex and black-box optimization problems. It consists in running a non-homogeneous Markov chain to sample from a sequence of Boltzmann probability distributions. This sequence is controlled by a cooling schedule, which governs the concentration of the mass of the Boltzmann distributions around the global minimizers. However, convergence is often slow, difficult to assess, and requires a fixed cooling schedule. We propose here a new simulated annealing algorithm with adaptive cooling schedule, which draws samples from variational approximations of the Boltzmann distributions. Our approach is theoretically sound and relies on an alternating Bregman proximal-gradient scheme minimizing a regularized Rényi divergence. Numerical experiments illustrate the performance of the method.
Thomas Guilmeau, Emilie Chouzenoux, Victor Elvira
ICASSP3
2023 Adaptive Gaussian Nested Filter for Parameter Estimation and State Tracking in Dynamical Systems
abstract
We introduce the adaptive Gaussian nested filter (AGNesF), the first nested method that adapts the number of samples to estimate both the static parameters and the dynamical variables of a state-space model. The proposed method is based on the nested Gaussian filter (NGF), that combines two layers of inference, one inside the other, to compute the joint posterior probability distribution of the static parameters and the state variables. We propose two novel rules to reduce computational complexity without compromising the performance. One enables the bottom layer techniques to run recursively, while the other reduces automatically the number of samples in the parameter space when they are redundant. We describe a specific implementation of the new scheme that uses a quadrature Kalman filter (QKF) in the parameter layer, and we study its performance in a stochastic Lorenz 63 model.
Sara Pérez-Vieites, Victor Elvira
ICASSP2
2023 An Augmented Gaussian Sum Filter through a mixture Decomposition
abstract
Bayesian filtering is an approach to the inference problem for state-space models (SSM), arising in many disciplines of science and engineering. The well known Gaussian filters tackle the problem using Gaussian approximations of the distributions of the hidden state in the model. Gaussian filters however are unable to track multimodal distributions that commonly arise in complex dynamical systems. The Gaussian sum filter (GSF) addresses this problem by using Gaussian mixture approximations of the filtering distribution. The GSF however has important limitations, requiring small width of the component covariances for good performance. Moreover, in many SSMs the estimates provided by the GSF blow up, due to covariance inflation. In this paper, we propose a way of controlling the covariances of the underlying Gaussian mixture. Our approach relies on a well known Gaussian identity, which helps us break down each component of the GSF, the parent, into several children components of smaller width. These smaller components are propagated through the nonlinearities by local linearization, which results in a smaller error than in the standard GSF. To reduce the resulting mixture we use resampling. We refer to our novel approach as augmented Gaussian sum filter (AGSF). We demonstrate the advantages of our approach using a toy example, for which the extended Kalman filter (EKF) and GSF perform poorly due to covariance inflation.
Kostas Tsampourakis, Victor Elvira
ICASSP2
2023 Cooperation in the Latent Space: The Benefits of Adding Mixture Components in Variational Autoencoders
abstract
In this paper, we show how the mixture components cooperate when they jointly adapt to maximize the ELBO. We build upon recent advances in the multiple and adaptive importance sampling literature. We then model the mixture components using separate encoder networks and show empirically that the ELBO is monotonically non-decreasing as a function of the number of mixture components. These results hold for a range of different VAE architectures on the MNIST, FashionMNIST, and CIFAR-10 datasets. In this work, we also demonstrate that increasing the number of mixture components improves the latent-representation capabilities of the VAE on both image and single-cell datasets. This cooperative behavior motivates that using Mixture VAEs should be considered a standard approach for obtaining more flexible variational approximations. Finally, Mixture VAEs are here, for the first time, compared and combined with normalizing flows, hierarchical models and/or the VampPrior in an extensive ablation study. Multiple of our Mixture VAEs achieve state-of-the-art log-likelihood results for VAE architectures on the MNIST and FashionMNIST datasets. The experiments are reproducible using our code, provided https://github.com/Lagergren-Lab/MixtureVAEs.
Oskar Kviman, Ricky Molén, Alexandra Hotti, Semih Kurt, Victor Elvira, Jens Lagergren
ICML5
2022 Multiple Importance Sampling ELBO and Deep Ensembles of Variational Approximations
abstract
In variational inference (VI), the marginal log-likelihood is estimated using the standard evidence lower bound (ELBO), or improved versions as the importance weighted ELBO (IWELBO). We propose the multiple importance sampling ELBO (MISELBO), a versatile yet simple framework. MISELBO is applicable in both amortized and classical VI, and it uses ensembles, e.g., deep ensembles, of independently inferred variational approximations. As far as we are aware, the concept of deep ensembles in amortized VI has not previously been established. We prove that MISELBO provides a tighter bound than the average of standard ELBOs, and demonstrate empirically that it gives tighter bounds than the average of IWELBOs. MISELBO is evaluated in density-estimation experiments that include MNIST and several real-data phylogenetic tree inference problems. First, on the MNIST dataset, MISELBO boosts the density-estimation performances of a state-of-the-art model, nouveau VAE. Second, in the phylogenetic tree inference setting, our framework enhances a state-of-the-art VI algorithm that uses normalizing flows. On top of the technical benefits of MISELBO, it allows to unveil connections between VI and recent advances in the importance sampling literature, paving the way for further methodological advances. We provide our code at https://github.com/Lagergren-Lab/MISELBO.
Oskar Kviman, Harald Melin, Hazal Koptagel, Victor Elvira, Jens Lagergren
AISTATS4
2022 Proximal-Based Adaptive Simulated Annealing for Global Optimization
abstract
Simulated annealing (SA) is a widely used approach to solve global optimization problems in signal processing. The initial non-convex problem is recast as the exploration of a sequence of Boltzmann probability distributions, which are increasingly harder to sample from. They are parametrized by a temperature that is iteratively decreased, following the so-called cooling schedule. Convergence results of SA methods usually require the cooling schedule to be set a priori with slow decay. In this work, we introduce a new SA approach that selects the cooling schedule on the fly. To do so, each Boltzmann distribution is approximated by a proposal density, which is also sequentially adapted. Starting from a variational formulation of the problem of joint temperature and proposal adaptation, we derive an alternating Bregman proximal algorithm to minimize the resulting cost, obtaining the sequence of Boltzmann distributions and proposals. Numerical experiments in an idealized setting illustrate the potential of our method compared with state-of-the-art SA algorithms.
Thomas Guilmeau, Emilie Chouzenoux, Victor Elvira
ICASSP3
2022 Approximating The Likelihood Ratio in Linear-Gaussian State-Space Models for Change Detection
abstract
Change-point detection methods are widely used in signal processing, primarily for detecting and locating changes in a considered model. An important family of algorithms for this problem relies on the likelihood ratio (LR) test. In state-space models (SSMs), the time series is modeled through a Markovian latent process. In this paper, we focus on the linear-Gaussian (LG) SSM, in which the LR-based methods require running a Kalman filter for every candidate change point. Since the number of candidates grows with the length of the time series, this strategy is inefficient in short time series and unfeasible for long ones. We propose an approximation to the LR which uses a constant number of filters, independently on the time-series length. The approximated LR relies on the Markovian property of the filter, which forgets errors at an exponential rate. We present theoretical results that justify the approximation, and we bound its error. We demonstrate its good performance in two numerical examples.
Kostas Tsampourakis, Victor Elvira
ICASSP2
2021 Comparison of Discrete and Continuous State Estimation with Focus on Active Flux Scheme
Jakub Matousek, Jindrich Duník, Marek Brandner, Victor Elvira
FUSION4
2021 Importance Gauss-Hermite Gaussian Filter for Models with Non-Additive Non-Gaussian Noises
Ondrej Straka, Jindrich Duník, Victor Elvira
FUSION3
2021 Would Your Tweet Invoke Hate on the Fly? Forecasting Hate Intensity of Reply Threads on Twitter
abstract
Curbing hate speech is undoubtedly a major challenge for online microblogging platforms like Twitter. While there have been studies around hate speech detection, it is not clear how hate speech finds its way into an online discussion. It is important for a content moderator to not only identify which tweet is hateful but also to predict which tweet will be responsible for accumulating hate speech. This would help in prioritizing tweets that need constant monitoring. Our analysis reveals that for hate speech to manifest in an ongoing discussion, the source tweet may not necessarily be hateful; rather, there are plenty of such non-hateful tweets which gradually invoke hateful replies, resulting in the entire reply threads becoming provocative.
Snehil Dahiya, Dhruv Sahnan, Vasu Goel, Emilie Chouzenoux, Victor Elvira, Angshul Majumdar, Anil Bandhakavi, Tanmoy Chakraborty 0002
KDD6
2021 Optimized auxiliary particle filters: adapting mixture proposals via convex optimization
abstract
Auxiliary particle filters (APFs) are a class of sequential Monte Carlo (SMC) methods for Bayesian inference in state-space models. In their original derivation, APFs operate in an extended state space using an auxiliary variable to improve inference. In this work, we propose optimized auxiliary particle filters, a framework where the traditional APF auxiliary variables are interpreted as weights in a importance sampling mixture proposal. Under this interpretation, we devise a mechanism for proposing the mixture weights that is inspired by recent advances in multiple and adaptive importance sampling. In particular, we propose to select the mixture weights by formulating a convex optimization problem, with the aim of approximating the filtering posterior at each timestep. Further, we propose a weighting scheme that generalizes previous results on the APF (Pitt et al. 2012), proving unbiasedness and consistency of our estimators. Our framework demonstrates significantly improved estimates on a range of metrics compared to state-of-the-art particle filters at similar computational complexity in challenging and widely used dynamical models.
Nicola Branchini, Victor Elvira
UAI2
2021 Recurrent dictionary learning for state-space models with an application in stock forecasting
Victor Elvira, Emilie Chouzenoux, Angshul Majumdar
Neurocomputing2
2021 Compressed Monte Carlo with application in particle filtering
Luca Martino, Victor Elvira
Inf. Sci.2
2021 Hamiltonian Adaptive Importance Sampling
abstract
Importance sampling (IS) is a powerful Monte Carlo (MC) methodology for approximating integrals, for instance in the context of Bayesian inference. In IS, the samples are simulated from the so-called proposal distribution, and the choice of this proposal is key for achieving a high performance. In adaptive IS (AIS) methods, a set of proposals is iteratively improved. AIS is a relevant and timely methodology although many limitations remain yet to be overcome, e.g., the curse of dimensionality in high-dimensional and multi-modal problems. Moreover, the Hamiltonian Monte Carlo (HMC) algorithm has become increasingly popular in machine learning and statistics. HMC has several appealing features such as its exploratory behavior, especially in high-dimensional targets, when other methods suffer. In this letter, we introduce the novel Hamiltonian adaptive importance sampling (HAIS) method. HAIS implements a two-step adaptive process with parallel HMC chains that cooperate at each iteration. The proposed HAIS efficiently adapts a population of proposals, extracting the advantages of HMC. HAIS can be understood as a particular instance of the generic layered AIS family with an additional resampling step. HAIS achieves a significant performance improvement in high-dimensional problems w.r.t. state-of-the-art algorithms. We discuss the statistical properties of HAIS and show its high performance in two challenging examples.
Reza Monsefi, Victor Elvira
IEEE Signal Process. Lett.3
2021 An Efficient Sampling Scheme for the Eigenvalues of Dual Wishart Matrices
abstract
Despite the numerous results in the literature about the eigenvalue distributions of Wishart matrices, the existing closed-form probability density function (pdf) expressions do not allow for efficient sampling schemes from such densities. In this letter, we present a stochastic representation for the eigenvalues of$2 \times 2$complex central uncorrelated Wishart matrices with an arbitrary number of degrees of freedom (referred to as dual Wishart matrices). The draws from the joint pdf of the eigenvalues are generated by means of a simple transformation of a chi-squared random variable and an independent beta random variable. Moreover, this stochastic representation allows a simple derivation, alternative to those already existing in the literature, of some eigenvalue function distributions such as the condition number or the ratio of the maximum eigenvalue to the trace of the matrix. The proposed sampling scheme may be of interest in wireless communications and multivariate statistical analysis, where Wishart matrices play a central role.
Ignacio Santamaría, Victor Elvira
IEEE Signal Process. Lett.2
2020 Graphem: EM Algorithm for Blind Kalman Filtering Under Graphical Sparsity Constraints
abstract
Modeling and inference with multivariate sequences is central in a number of signal processing applications such as acoustics, social network analysis, biomedical, and finance, to name a few. The linear-Gaussian state-space model is a common way to describe a time series through the evolution of a hidden state, with the advantage of presenting a simple inference procedure due to the celebrated Kalman filter. A fundamental question when analyzing a multivariate sequence is the search for relationships between its entries (or the entries of the modeled hidden state), especially when the inherent structure is a non-fully connected graph. In such context, graphical modeling combined with parsimony constraints allows to limit the proliferation of parameters and enables a compact data representation which is easier to interpret by the experts. In this work, we propose a novel expectation-maximization algorithm for estimating the linear matrix operator in the state equation of a linear-Gaussian state-space model. Lasso regularization is included in the M-step, that we solve using a proximal splitting Douglas-Rachford algorithm. Numerical experiments illustrate the benefits of the proposed model and inference technique, named GraphEM, over competitors relying on Granger causality.
Emilie Chouzenoux, Victor Elvira
ICASSP2
2020 Particle Group Metropolis Methods for Tracking the Leaf Area Index
abstract
Monte Carlo (MC) algorithms are widely used for Bayesian inference in statistics, signal processing, and machine learning. In this work, we introduce an Markov Chain Monte Carlo (MCMC) technique driven by a particle filter. The resulting scheme is a generalization of the so-called Particle Metropolis-Hastings (PMH) method, where a suitable Markov chain of sets of weighted samples is generated. We also introduce a marginal version for the goal of jointly inferring dynamic and static variables. The proposed algorithms outperform the corresponding standard PMH schemes, as shown by numerical experiments.
Luca Martino, Victor Elvira, Gustau Camps-Valls
ICASSP2
2019 Langevin-based Strategy for Efficient Proposal Adaptation in Population Monte Carlo
abstract
Population Monte Carlo (PMC) algorithms are a family of adaptive importance sampling (AIS) methods for approximating integrals in Bayesian inference. In this paper, we propose a novel PMC algorithm that combines recent advances in the AIS and the optimization literatures. In such a way, the proposal densities are adapted according to the past weighted samples via a local resampling that preserves the diversity, but we also exploit the geometry of the targeted distribution. A scaled Langevin strategy with Newton-based scaling metric is retained for this purpose, allowing to adapt jointly the means and the covariances of the proposals, without needing to tune any extra parameter. The performance of the proposed technique is clearly superior in two numerical examples at the cost of a reasonable computational complexity increment.
Victor Elvira, Emilie Chouzenoux
ICASSP1
2019 A Probabilistic Incremental Proximal Gradient Method
abstract
In this letter, we propose a probabilistic optimization method, named probabilistic incremental proximal gradient (PIPG) method, by developing a probabilistic interpretation of the incremental proximal gradient algorithm. We explicitly model the update rules of the incremental proximal gradient method and develop a systematic approach to propagate the uncertainty of the solution estimate over iterations. The PIPG algorithm takes the form of Bayesian filtering updates for a state-space model constructed by using the cost function. Our framework makes it possible to utilize well-known exact or approximate Bayesian filters, such as Kalman or extended Kalman filters, to solve large-scale regularized optimization problems.
Ömer Deniz Akyildiz, Emilie Chouzenoux, Victor Elvira, Joaquín Míguez
IEEE Signal Process. Lett.3
2019 Multiple Importance Sampling for Efficient Symbol Error Rate Estimation
abstract
Digital constellations formed by hexagonal or other non-square two-dimensional lattices are often used in advanced digital communication systems. The integrals required to evaluate the symbol error rate (SER) of these constellations in the presence of Gaussian noise are in general difficult to compute in closed form, and therefore Monte Carlo simulation is typically used to estimate the SER. However, naive Monte Carlo simulation can be very inefficient and requires very long simulation runs, especially at high signal-to-noise ratios. In this letter, we adapt a recently proposed multiple importance sampling technique, called ALOE (for “at least one rare event”), to this problem. Conditioned to a transmitted symbol, an error (or rare event) occurs when the observation falls in a union of half-spaces or, equivalently, outside a given polytope. The proposal distribution for ALOE samples the system conditionally on an error taking place, which makes it more efficient than other importance sampling techniques. ALOE provides unbiased SER estimates with simulation times orders of magnitude shorter than conventional Monte Carlo.
Victor Elvira, Ignacio Santamaría
IEEE Signal Process. Lett.1
2019 Hierarchical Algorithms for Causality Retrieval in Atrial Fibrillation Intracavitary Electrograms
abstract
Multichannel intracavitary electrograms (EGMs) are acquired at the electrophysiology laboratory to guide radio frequency catheter ablation of patients suffering from atrial fibrillation. These EGMs are used by cardiologists to determine candidate areas for ablation (e.g., areas corresponding to high dominant frequencies or complex fractionated electrograms). In this paper, we introduce two hierarchical algorithms to retrieve the causal interactions among these multiple EGMs. Both algorithms are based on Granger causality, but other causality measures can be easily incorporated. In both cases, they start by selecting a root node, but they differ on the way in which they explore the set of signals to determine their cause-effect relationships: either testing the full set of unexplored signals (GS-CaRe) or performing a local search only among the set of neighbor EGMs (LS-CaRe). The ensuing causal model provides important information about the propagation of the electrical signals inside the atria, uncovering wavefronts and activation patterns that can guide cardiologists towards candidate areas for catheter ablation. Numerical experiments, on both synthetic signals and annotated real-world signals, show the good performance of the two proposed approaches.
David Luengo, Gonzalo R. Ríos-Muñoz, Victor Elvira, Carlos Sánchez 0005, Antonio Artés-Rodríguez
IEEE J. Biomed. Health Informatics3
2018 The Incremental Proximal Method: A Probabilistic Perspective
abstract
In this work, we highlight a connection between the incremental proximal method and stochastic filters. We begin by showing that the proximal operators coincide, and hence can be realized with, Bayes updates. We give the explicit form of the updates for the linear regression problem and show that there is a one-to-one correspondence between the proximal operator of the least-squares regression and the Bayes update when the prior and the likelihood are Gaussian. We then carry out this observation to a general sequential setting: We consider the incremental proximal method, which is an algorithm for large-scale optimization, and show that, for a linear-quadratic cost function, it can naturally be realized by the Kalman filter. We then discuss the implications of this idea for nonlinear optimization problems where proximal operators are in general not realizable. In such settings, we argue that the extended Kalman filter can provide a systematic way for the derivation of practical procedures.
Ömer Deniz Akyildiz, Victor Elvira, Joaquín Míguez
ICASSP2
2018 Pattern Localization in Time Series Through Signal-To-Model Alignment in Latent Space
abstract
In this paper, we study the problem of locating a predefined sequence of patterns in a time series. In particular, the studied scenario assumes a theoretical model is available that contains the expected locations of the patterns. This problem is found in several contexts, and it is commonly solved by first synthesizing a time series from the model, and then aligning it to the true time series through dynamic time warping. We propose a technique that increases the similarity of both time series before aligning them, by mapping them into a latent correlation space. The mapping is learned from the data through a machine-learning setup. Experiments on data from nondestructive testing demonstrate that the proposed approach shows significant improvements over the state of the art.
Steven Van Vaerenbergh, Ignacio Santamaría, Victor Elvira, Matteo Salvatori
ICASSP3
2018 Robust Covariance Adaptation in Adaptive Importance Sampling
abstract
Importance sampling (IS) is a Monte Carlo methodology that allows for the approximation of a target distribution using weighted samples generated from another proposal distribution. Adaptive importance sampling (AIS) implements an iterative version of IS, which adapts the parameters of the proposal distribution in order to improve estimation of the target. While the adaptation of the location (mean) of the proposals has been largely studied, an important challenge of AIS relates to the difficulty of adapting the scale parameter (covariance matrix). In the case of weight degeneracy, adapting the covariance matrix using the empirical covariance results in a singular matrix, which leads to a poor performance in subsequent iterations of the algorithm. In this letter, we propose a novel scheme which exploits recent advances in the IS literature to prevent the so-called weight degeneracy. The method efficiently adapts the covariance matrix of a population of proposal distributions and achieves a significant performance improvement in high-dimensional scenarios. We validate the new method through computer simulations.
Yousef El-Laham, Victor Elvira, Mónica F. Bugallo
IEEE Signal Process. Lett.2
2018 Multiple importance sampling characterization by weighted mean invariance
Mateu Sbert, Vlastimil Havran, László Szirmay-Kalos, Victor Elvira
Vis. Comput.4
2017 Improving population Monte Carlo: Alternative weighting and resampling schemes
Victor Elvira, Luca Martino, David Luengo, Mónica F. Bugallo
Signal Process.1
2017 Effective sample size for importance sampling based on discrepancy measures
Luca Martino, Victor Elvira, Francisco Louzada 0001
Signal Process.2
2016 Online adaptation of the number of particles of SMC methods
abstract
Particle filtering is a widely used sequential methodology that approximates probability distributions by using discrete random measures composed of weighted particles. A large number of particles improves the quality of the approximation but increases the computational requirements. Although there exists an abundant variety of particle filtering algorithms in the literature, there is lack of work devoted to selecting or adapting the number of particles systematically. In this paper we propose a novel methodology for online assessment of convergence of particle filtering. Based on theoretical analysis of the assessment, we propose an algorithm for the adaptation of the number of particles in online manner. The performance of the proposed algorithm is demonstrated for two state-space models.
Victor Elvira, Joaquín Míguez, Petar M. Djuric
ICASSP1
2016 A hierarchical algorithm for causality discovery among atrial fibrillation electrograms
abstract
Multi-channel intracardiac electrocardiograms (electrograms) are sequentially acquired, at the electrophysiology laboratory, in order to guide radio frequency catheter ablation during heart surgery performed on patients with sustained atrial fibrillation (AF). These electrograms are used by cardiologists to determine candidate areas for ablation (e.g., areas corresponding to high dominant frequencies or complex fractionated electrograms). In this paper, we introduce a novel hierarchical algorithm for causality discovery among these multi-output sequentially acquired electrograms. The causal model obtained provides important information about the propagation of the electrical signals inside the heart, uncovering wavefronts and activation patterns that will serve to increase our knowledge about AF and guide cardiologists towards candidate areas for catheter ablation. Numerical results on synthetic signals, generated using the FitzHugh-Nagumo model, show the good performance of the proposed approach.
David Luengo, Gonzalo R. Ríos-Muñoz, Victor Elvira, Antonio Artés-Rodríguez
ICASSP3
2016 Parallel metropolis chains with cooperative adaptation
abstract
Monte Carlo methods, such as Markov chain Monte Carlo (MCMC) algorithms, have become very popular in signal processing over the last years. In this work, we introduce a novel MCMC scheme where parallel MCMC chains interact, adapting cooperatively the parameters of their proposal functions. Furthermore, the novel algorithm distributes the computational effort adaptively, rewarding the chains which are providing better performance and, possibly even stopping other ones. These extinct chains can be reactivated if the algorithm considers it necessary. Numerical simulations show the benefits of the novel scheme.
Luca Martino, Victor Elvira, David Luengo, Francisco Louzada 0001
ICASSP2
2016 Heretical Multiple Importance Sampling
abstract
Multiple importance sampling (MIS) methods approximate moments of complicated distributions by drawing samples from a set of proposal distributions. Several ways to compute the importance weights assigned to each sample have been recently proposed, with the so-called deterministic mixture (DM) weights providing the best performance in terms of variance, at the expense of an increase in the computational cost. A recent work has shown that it is possible to achieve a tradeoff between variance reduction and computational effort by performing an a priori random clustering of the proposals (partial DM algorithm). In this paper, we propose a novel “heretical” MIS framework, where the clustering is performed a posteriori with the goal of reducing the variance of the importance sampling weights. This approach yields biased estimators with a potentially large reduction in variance. Numerical examples show that heretical MIS estimators can outperform, in terms of mean squared error, both the standard and the partial MIS estimators, achieving a performance close to that of DM with less computational cost.
Victor Elvira, Luca Martino, David Luengo, Mónica F. Bugallo
IEEE Signal Process. Lett.1
2015 A gradient adaptive population importance sampler
abstract
Monte Carlo (MC) methods are widely used in signal processing and machine learning. A well-known class of MC methods is composed of importance sampling and its adaptive extensions (e.g., population Monte Carlo). In this paper, we introduce an adaptive importance sampler using a population of proposal densities. The novel algorithm dynamically optimizes the cloud of proposals, adapting them using information about the gradient and Hessian matrix of the target distribution. Moreover, a new kind of interaction in the adaptation of the proposal densities is introduced, establishing a trade-off between attaining a good performance in terms of mean square error and robustness to initialization.
Victor Elvira, Luca Martino, David Luengo, Jukka Corander
ICASSP1
2015 A probabilistic least-mean-squares filter
abstract
We introduce a probabilistic approach to the LMS filter. By means of an efficient approximation, this approach provides an adaptable step-size LMS algorithm together with a measure of uncertainty about the estimation. In addition, the proposed approximation preserves the linear complexity of the standard LMS. Numerical results show the improved performance of the algorithm with respect to standard LMS and state-of-the-art algorithms with similar complexity. The goal of this work, therefore, is to open the door to bring somemore Bayesian machine learning techniques to adaptive filtering.
Jesus Fernandez-Bes, Victor Elvira, Steven Van Vaerenbergh
ICASSP2
2015 Efficient linear combination of partial Monte Carlo estimators
abstract
In many practical scenarios, including those dealing with large data sets, calculating global estimators of unknown variables of interest becomes unfeasible. A common solution is obtaining partial estimators and combining them to approximate the global one. In this paper, we focus on minimum mean squared error (MMSE) estimators, introducing two efficient linear schemes for the fusion of partial estimators. The proposed approaches are valid for any type of partial estimators, although in the simulated scenarios we concentrate on the combination of Monte Carlo estimators due to the nature of the problem addressed. Numerical results show the good performance of the novel fusion methods with only a fraction of the cost of the asymptotically optimal solution.
David Luengo, Luca Martino, Victor Elvira, Mónica F. Bugallo
ICASSP3
2015 Smelly parallel MCMC chains
abstract
Monte Carlo (MC) methods are useful tools for Bayesian inference and stochastic optimization that have been widely applied in signal processing and machine learning. A well-known class of MC methods are Markov Chain Monte Carlo (MCMC) algorithms. In this work, we introduce a novel parallel interacting MCMC scheme, where the parallel chains share information, thus yielding a faster exploration of the state space. The interaction is carried out generating a dynamic repulsion among the “smelly” parallel chains that takes into account the entire population of current states. The ergodicity of the scheme and its relationship with other sampling methods are discussed. Numerical results show the advantages of the proposed approach in terms of mean square error, robustness w.r.t. to initial values and parameter choice.
Luca Martino, Victor Elvira, David Luengo, Antonio Artés-Rodríguez, Jukka Corander
ICASSP2
2015 Efficient Multiple Importance Sampling Estimators
abstract
Multiple importance sampling (MIS) methods use a set of proposal distributions from which samples are drawn. Each sample is then assigned an importance weight that can be obtained according to different strategies. This work is motivated by the trade-off between variance reduction and computational complexity of the different approaches (classical vs. deterministic mixture) available for the weight calculation. A new method that achieves an efficient compromise between both factors is introduced in this letter. It is based on forming a partition of the set of proposal distributions and computing the weights accordingly. Computer simulations show the excellent performance of the associated partial deterministic mixture MIS estimator.
Victor Elvira, Luca Martino, David Luengo, Mónica F. Bugallo
IEEE Signal Process. Lett.1
2014 An adaptive population importance sampler
abstract
Monte Carlo (MC) methods are widely used in signal processing, machine learning and communications for statistical inference and stochastic optimization. A well-known class of MC methods is composed of importance sampling and its adaptive extensions (e.g., population Monte Carlo). In this work, we introduce an adaptive importance sampler using a population of proposal densities. The novel algorithm provides a global estimation of the variables of interest iteratively, using all the samples generated. The cloud of proposals is adapted by learning from a subset of previously generated samples, in such a way that local features of the target density can be better taken into account compared to single global adaptation procedures. Numerical results show the advantages of the proposed sampling scheme in terms of mean absolute error and robustness to initialization.
Luca Martino, Victor Elvira, David Luengo, Jukka Corander
ICASSP2
2012 Analog antenna combining in transmit correlated channels: Transceiver design and performance evaluation
Victor Elvira, Javier Vía
Signal Process.1
2010 A General Pre-FFT Criterion for MIMO-OFDM Beamforming
abstract
In this paper, we propose a general beamforming criterion for pre-FFT processing in orthogonal frequency division multiplexing (OFDM) systems with multiple transmit and receive antennas. The proposed criterion depends on a single parameter α, which establishes a tradeoff between the energy of the equivalent SISO channel (after Tx-Rx beamforming) and its spectral flatness. The proposed cost function embraces most reasonable criteria for designing Tx-Rx pre-FFT beamformers. Hence, for particular values of α the proposed criterion reduces to the minimization of the mean square error (MSE), the maximization of the system capacity, or the maximization of the received signal-to-noise ratio (SNR). In general, the proposed criterion results in a non convex optimization problem. However, we show that the problem can be approximately solved by semidefinite relaxation (SDR) techniques. Additionally, since the computational cost of SDR for this problem is rather high, we propose a simple yet efficient gradient search algorithm which provides satisfactory solutions with a moderate computational cost for OFDM-based WLAN standards such as 802.11a. Finally, the good performance of the proposed technique is illustrated by means of some numerical results.
Javier Vía, Ignacio Santamaría, Victor Elvira, Ralf Eickhoff
ICC3
2009 Minimum BER beamforming in the RF domain for OFDM transmissions and linear receivers
abstract
In this paper, we study transmission schemes for a novel OFDM-based MIMO system which performs adaptive signal combining in radio-frequency (RF). Specifically, we consider the problem of selecting the linear precoder and the transmit and receive RF weights (or beamformers) for minimizing the bit error rate (BER) under the assumption of perfect channel knowledge and linear receivers. Firstly, it is shown that the optimal precoder amounts to uniformly distribute the overall mean square error (MSE) among the information symbols. Secondly, we propose a gradient search algorithm to obtain the optimal pair of beamformers. Interestingly, in the case of low signal to noise ratios (SNR), the proposed beamforming criterion is equivalent to the maximization of the received SNR. However, for moderate and high SNRs, part of the received SNR is sacrificed in order to improve the channel response of the worst subcarriers, which translates into significant advantages over other previously proposed approaches. Finally, the performance of the proposed scheme is illustrated by means of some numerical examples.
Javier Vía, Victor Elvira, Ignacio Santamaría, Ralf Eickhoff
ICASSP2
2009 Analog Antenna Combining for Maximum Capacity Under OFDM Transmissions
abstract
In this paper, we study beamforming schemes for a novel MIMO transceiver, which performs adaptive signal combining in the radio-frequency domain. Assuming perfect channel knowledge at both the transmit and receive sides, we consider the problem of selecting the transmit and receive RF beamformers that maximize the capacity (MaxCAP criterion) of the system under orthogonal frequency division multiplexing (OFDM) transmissions. This problem is non-convex and has no closed-form solution, therefore the maximum capacity beamformers are found using a gradient search algorithm. Furthermore, it is shown in the paper that, for low signal-to-noise ratios (SNR), the MaxCAP criterion is equivalent to maximizing the received SNR (MaxSNR criterion). However, for moderate and high SNRs, the maximum capacity beamformers sacrifice part of the received SNR in order to improve the worst subcarriers and, in this way, they increase the overall capacity of the multicarrier channel. Finally, by means of numerical examples we show that the MaxCAP criterion significantly outperforms the MaxSNR criterion in terms of bit error rate and outage probability.
Javier Vía, Victor Elvira, Ignacio Santamaría, Ralf Eickhoff
ICC2