Filip Elvander

dblp:173/6871 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0003-1857-2173ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Sound Field Estimation Using Optimal Transport Barycenters in the Presence of Phase Errors
abstract
Publisher Copyright: © 1994-2012 IEEE.
Filip Elvander
IEEE Signal Process. Lett.3
2025 Robust Multi-Pitch Estimation via Optimal Transport Clustering
abstract
In this work, we consider the multi-pitch estimation problem, i.e., to estimate multiple sets of harmonically related sinusoids from noisy measurements. We propose to phrase this as a clustering problem with indirect measurements, where we simultaneously infer the spectral content of the signal and group its power into a small set of harmonic structures. The grouping is enforced using a regularization function building on optimal transport theory. The resulting estimator is formulated in terms of the solution of a convex optimization problem, and we present an efficient algorithm implementing the estimator. In numerical experiments, we show that the proposed estimator displays competitive performance as compared to the state-of-the-art. In particular, the proposed estimator is shown to be highly robust to inharmonicities, i.e., deviations from perfect harmonicity.
Anton Björkman, Filip Elvander
ICASSP2
2024 Estimation of Impulse Responses for a Moving Source Using Optimal Transport Regularization
abstract
The estimation of impulse responses (IRs) is fundamental to various audio applications, including active noise control, telecommunication, and sound zone control. Despite its long history, estimating impulse responses remains challenging when dealing with short signals and with signals having poor spectral excitation. However, in many applications the source is moving such that one has access to several input-output signal pairs corresponding to closely spaced source positions. Intuitively, exploiting this spatial proximity when jointly estimating the full set of IRs should allow for improved estimation performance. In this work, we propose to leverage the information shared between the closely spaced source positions by means of an optimal transport regularizer when estimating IRs from noisy input-output relations. In particular, the proposed transport formulation allows for modeling shifts in time-delays, corresponding to the IR filter taps, caused by the spatial displacement. The method is validated through numerical experiments using a real voice recording as input signal, demonstrating its superior performance in the challenging scenario.
David Sundström, Filip Elvander, Andreas Jakobsson
ICASSP2
2024 Blind Audio Bandwidth Extension: A Diffusion-Based Zero-Shot Approach
abstract
Audio bandwidth extension involves the realistic reconstruction of high-frequency spectra from bandlimited observations. In cases where the lowpass degradation is unknown, such as in restoring historical audio recordings, this becomes a blind problem. This paper introduces a novel method called BABE (Blind Audio Bandwidth Extension) that addresses the blind problem in a zero-shot setting, leveraging the generative priors of a pre-trained unconditional diffusion model. During the inference process, BABE utilizes a generalized version of diffusion posterior sampling, where the degradation operator is unknown but parametrized and inferred iteratively. The performance of the proposed method is evaluated using objective and subjective metrics, and the results show that BABE surpasses state-of-the-art blind bandwidth extension baselines and achieves competitive performance compared to informed methods when tested with synthetic data. Moreover, BABE exhibits robust generalization capabilities when enhancing real historical recordings, effectively reconstructing the missing high-frequency content while maintaining coherence with the original recording. Subjective preference tests confirm that BABE significantly improves the audio quality of historical music recordings. Examples of historical recordings restored with the proposed method are available on the companion webpage:http://research.spa.aalto.fi/publications/papers/ieee-taslp-babe/
Eloi Moliner, Filip Elvander, Vesa Välimäki
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Multi-Source Direction-of-Arrival Estimation Using Steered Response Power and Group-Sparse Optimization
abstract
In this paper, a method is proposed for estimating the direction of arrival (DOA) of multiple broadband sound sources. This is achieved by solving a group-sparse optimization problem, which models an observed broadband steered response power (SRP) map as a linear function of power spectral densities (PSDs), associated to a set of candidate DOAs. The estimation of the source DOAs is then accomplished by identifying peaks in the resulting spatial power density, i.e., the estimated direction-specific PSDs integrated over frequency. The proposed method is motivated by its potential to reveal more distinct peaks in the estimated spatial power density than those directly observed in the SRP map, which can be beneficial to the robustness in DOA estimation performance when multiple sources need to be distinguished under varying acoustic conditions. An implementation of the proposed method using the alternating direction method of multipliers (ADMM) is presented, and the DOA estimation performance is evaluated with both simulated and experimental data. Results show that, especially in reverberant scenarios, the proposed method presents an advantage in locating closely spaced sources when compared to SRP-PHAT, the group-sparse iterative covariance-based estimation (GSPICE) method, and the wideband MUSIC method with geometric averaging. Furthermore, it is observed that for a compact microphone array, the proposed method overall maintained its performance even when using SRP maps with lower grid resolutions than the sampling requirements of the broadband SRP function. Finally, results obtained with experimental data showed the applicability of the proposed method in a practical meeting room environment.
Elisa Tengan, Thomas Dietzen, Filip Elvander, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Distributed Adaptive Norm Estimation for Blind System Identification in Wireless Sensor Networks
abstract
Distributed signal-processing algorithms in (wireless) sensor networks often aim to decentralize processing tasks to reduce communication cost and computational complexity or avoid reliance on a single device (i.e., fusion center) for processing. In this contribution, we extend a distributed adaptive algorithm for blind system identification that relies on the estimation of a stacked network-wide consensus vector at each node, the computation of which requires either broadcasting or relaying of node-specific values (i.e., local vector norms) to all other nodes. The extended algorithm employs a distributed-averaging-based scheme to estimate the network-wide consensus norm value by only using the local vector norm provided by neighboring sensor nodes. We introduce an adaptive mixing factor between instantaneous and recursive estimates of these norms for adaptivity in a time-varying system. Simulation results show that the extension provides estimation results close to the optimal fully-connected-network or broadcasting case while reducing inter-node transmission significantly.
Matthias Blochberger, Filip Elvander, Randall Ali, Jan Østergaard, Jesper Jensen 0001, Marc Moonen, Toon van Waterschoot
ICASSP2
2023 Estimating Inharmonic Signals with Optimal Transport Priors
abstract
In this work, we consider the problem of estimating the frequency content of inharmonic signals, i.e., sinusoidal mixtures whose components are close to forming a harmonic set. Intuitively, exploiting this closeness should lead to increased estimation performance as compared to unstructured estimation. Earlier approaches to this problem have relied on parametric descriptions of the inharmonicity, stochastic representations, or have resorted to misspecified estimation by ignoring the inharmonicity. Herein, we propose to use a penalized maximum-likelihood framework, where the regularizer is constructed based on optimal mass transport theory, promoting estimates that are close-to-harmonic in a spectral sense. This leads to an estimator that forms a smooth path between the unstructured maximum-likelihood estimator (MLE) and a misspecified MLE (MMLE), as determined by a regularization parameter. In numerical illustrations, we show that the proposed estimator worst-case dominates the MLE and MMLE, thereby allowing for robust estimation for cases when the inharmonicity level is unknown.
Filip Elvander
ICASSP1
2023 Fast Low-Latency Convolution by Low-Rank Tensor Approximation
abstract
In this paper we consider fast time-domain convolution, exploiting low-rank properties of an impulse response (IR). This reduces the computational complexity, speeding up the convolution, without introducing latency. Previous work has considered a truncated singular value decomposition (SVD) of a two-dimensional matricization, or reshaping, of the IR. We here build upon this idea, by providing an algorithm for convolution with a three-dimensional tensorization of the IR. We provide simulations using real-life acoustic room impulse responses (RIRs) of various lengths, convolving them with music, as well as speech signals. The proposed algorithm is shown to outperform the comparable existing algorithm in terms of signal quality degradation, for all considered scenarios, without increasing the computational complexity, or the memory usage.
Martin Jälmby, Filip Elvander, Toon van Waterschoot
ICASSP2
2023 Simultaneous Acoustic Echo Sorting and 3-D Room Geometry Inference
abstract
Room geometry estimation from multiple acoustic room impulse responses (RIRs) relies on being able to correctly identify first-order echoes corresponding to the same reflector. This can be done by exploiting the properties of various mathematical concepts but often results in needing to solve large combinatorial problems owing to the increased number of measurements these tools require for efficacy. A low-complexity iterative method is proposed for determining the boundaries in convex polygonal rooms using common tangent planes to sets of ellipsoids. Prior knowledge of the order of echoes is unnecessary and only three RIRs are required, which is one less than the minimum number typically needed to find a unique reflector in 3-D, allowing for simpler common tangent plane estimation and reducing possible echo combinations. Candidate partial rooms are built by checking whether expected second-order echoes appear in the RIRs, adding a reflector at each iterate until a room is found. The proposed method is validated by means of computer simulations.
Kathleen MacWilliam, Filip Elvander, Toon van Waterschoot
ICASSP2
2023 Determining joint periodicities in multi-time data with sampling uncertainties
David Svedberg, Filip Elvander, Andreas Jakobsson
Signal Process.2
2023 Low-Rank Room Impulse Response Estimation
abstract
In this paper we consider low-rank estimation of room impulse responses (RIRs). Inspired by a physics-driven room-acoustical model, we propose an estimator of RIRs that promotes a low-rank structure for a matricization, or reshaping, of the estimated RIR. This low-rank prior acts as a regularizer for the inverse problem of estimating an RIR from input-output observations, preventing overfitting and improving estimation accuracy. As directly enforcing a low rank of the estimate results is an NP-hard problem, we consider two different relaxations, one using the nuclear norm, and one using the recently introduced concept of quadratic envelopes. Both relaxations allow for implementing the proposed estimator using a first-order algorithm with convergence guarantees. When evaluated on both synthetic and recorded RIRs, it is shown that under noisy output conditions, or when the spectral excitation of the input signal is poor, the proposed estimator outperforms comparable existing methods. The performance of the two low-rank relaxations methods is similar, but the quadratic envelope has the benefit of superior robustness to the choice of regularization hyperparameter in the case when the signal-to-noise ratio is unknown. The performance of the proposed method is compared to that of ordinary least squares, Tikhonov least squares, as well as the Cramér-Rao lower bound (CRLB).
Martin Jälmby, Filip Elvander, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Determining Joint Periodicities in Multi-Time Data with Sampling Uncertainties
abstract
In this work, we introduce a novel approach for determining a joint sparse spectrum from several non-uniformly sampled data sets, where each data set is assumed to have its own, and only partially known, sampling times. The problem originates in paleoclimatology, where each data point derives from a separate ice core measurement, resulting in that even though all measurements reflect the same periodicities, the sampling times and phases differ among the data sets, with the sampling times being only approximately known. The proposed estimator exploits all available data using a sparse reconstruction framework allowing for a reliable and robust estimation of the underlying periodicities. The performance of the method is illustrated using both simulated and measured ice core data sets.
David Svedberg, Filip Elvander, Andreas Jakobsson
ICASSP2
2020 On Harmonic Approximations of Inharmonic Signals
abstract
In this work, we present the misspecified Gaussian Cramér-Rao lower bound for the parameters of a harmonic signal, or pitch, when signal measurements are collected from an almost, but not quite, harmonic model. For the asymptotic case of large sample sizes, we present a closed-form expression for the bound corresponding to the pseduo-true fundamental frequency. Using simulation studies, it is shown that the bound is sharp and is attained by maximum likelihood estimators derived under the misspecified harmonic assumption. It is shown that misspecified harmonic models achieve a lower mean squared error than correctly specified unstructured models for moderately inharmonic signals. Examining voices from a speech database, we conclude that human speech belongs to this class of signals, verifying that the use of a harmonic model for voiced speech is preferable.
Filip Elvander, Andreas Jakobsson
ICASSP1
2020 Multi-marginal optimal transport using partial information with applications in robust localization and sensor fusion
Filip Elvander, Isabel Haasler, Andreas Jakobsson
Signal Process.1
2019 Non-coherent Sensor Fusion via Entropy Regularized Optimal Mass Transport
abstract
This work presents a method for information fusion in source localization applications. The method utilizes the concept of optimal mass transport in order to construct estimates of the spatial spectrum using a convex barycenter formulation. We introduce an entropy regularization term to the convex objective, which allows for low-complexity iterations of the solution algorithm and thus makes the proposed method applicable also to higher-dimensional problems. We illustrate the proposed method's inherent robustness to misalignment and miscalibration of the sensor arrays using numerical examples of localization in two dimensions.
Filip Elvander, Isabel Haasler, Andreas Jakobsson
ICASSP1
2018 Using Optimal Mass Transport for Tracking and Interpolation of Toeplitz Covariance Matrices
abstract
In this work, we propose a novel method for interpolation and extrapolation of Toeplitz structured covariance matrices. By considering a spectral representation of Toeplitz matrices, we use an optimal mass transport problem in the spectral domain in order to define a notion of distance between such matrices. The obtained optimal transport plan naturally induces a way of interpolating, as well as extrapolating, Toeplitz matrices. The constructed covariance matrix interpolants and ex-trapolants preserve the Toeplitz structure, as well as the positive semi-definiteness and the zeroth covariance of the original matrices. We demonstrate the proposed method's ability to model locally linear shifts of spectral power for slowly varying stochastic processes, illustrating the achievable performance using a simple tracking problem.
Filip Elvander, Andreas Jakobsson
ICASSP1
2018 Multi-dimensional grid-less estimation of saturated signals
Filip Elvander, Johan Sward, Andreas Jakobsson
Signal Process.1
2018 Designing sampling schemes for multi-dimensional data
Johan Sward, Filip Elvander, Andreas Jakobsson
Signal Process.2
2017 Using optimal transport for estimating inharmonic pitch signals
abstract
In this work, we propose a novel multi-pitch estimation technique that is robust with respect to the inharmonicity commonly occurring in many applications. The method does not require any a priori knowledge of the number of signal sources, the number of harmonics of each source, nor the structure or scope of any possibly occurring inharmonicity. Formulated as a minimum transport distance problem, the proposed method finds an estimate of the present pitches by mapping any found spectral line to the closest harmonic structure. The resulting optimization is a convex and highly tractable linear programming problem. The preferable performance of the proposed method is illustrated using both simulated and real audio signals.
Filip Elvander, Stefan Ingi Adalbjornsson, Andreas Jakobsson
ICASSP1
2017 Online Estimation of Multiple Harmonic Signals
abstract
In this paper, we propose a time-recursive multipitch estimation algorithm using a sparse reconstruction framework, assuming that only a few pitches from a large set of candidates are active at each time instant. The proposed algorithm does not require any training data, and instead utilizes a sparse recursive least-squares formulation augmented by an adaptive penalty term specifically designed to enforce a pitch structure on the solution. The amplitudes of the active pitches are also recursively updated, allowing for a smooth and more accurate representation. When evaluated on a set of ten music pieces, the proposed method is shown to outperform other general purpose multipitch estimators in either accuracy or computational speed, although not being able to yield performance as good as the state-of-the art methods, which are being optimally tuned and specifically trained on the present instruments. However, the method is able to outperform such a technique when used without optimal tuning, or when applied to instruments not included in the training data.
Filip Elvander, Johan Sward, Andreas Jakobsson
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 An adaptive penalty multi-pitch estimator with self-regularization
Filip Elvander, Ted Kronvall, Stefan Ingi Adalbjornsson, Andreas Jakobsson
Signal Process.1