Robin Scheibler

dblp:05/8655 · DBLP profile ↗
← Back
35ranked-venue papers
15as first author
20since 2021 · last 2025
0000-0002-5205-8365ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 14 first-author · 20 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 11 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Less is More: Data Curation Matters in Scaling Speech Enhancement
abstract
The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling speech enhancement data. We focus on a critical factor: prevalent quality issues in “clean” training labels within large-scale datasets. This work re-examines this phenomenon and demonstrates that, within large-scale training sets, prioritizing high-quality training data is more important than merely expanding the data volume. Experimental findings suggest that models trained on a carefully curated subset of 700 hours can outperform models trained on the 2,500 -hour full dataset. This outcome highlights the crucial role of data curation in scaling speech enhancement systems effectively.
Chenda Li, Wangyou Zhang, Wei Wang 0010, Robin Scheibler, Kohei Saijo, Samuele Cornell, Yihui Fu, Marvin Sach, Zhaoheng Ni, Anurag Kumar 0003, Tim Fingscheidt, Shinji Watanabe 0001, Yanmin Qian
ASRU4
2025 URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
abstract
The Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, such as DNSMOS and UTMOS, have been developed to predict MOS to avoid this issue, they often suffer from insufficient training data. Recognizing that the comparison of speech enhancement (SE) systems prioritizes a reliable system comparison over absolute scores, we propose URGENT-PK, a novel ranking approach leveraging pairwise comparisons. URGENT-PK takes homologous enhanced speech pairs as input to predict relative quality rankings. This pairwise paradigm efficiently utilizes limited training data, as all pairwise permutations of multiple systems constitute a training instance. Experiments across multiple open test sets demonstrate URGENT-PK’s superior system-level ranking performance over state-of-the-art baselines, despite its simple network architecture and limited training data.
Chenda Li, Wei Wang 0010, Wangyou Zhang, Samuele Cornell, Marvin Sach, Robin Scheibler, Kohei Saijo, Yihui Fu, Zhaoheng Ni, Anurag Kumar 0003, Tim Fingscheidt, Shinji Watanabe 0001, Yanmin Qian
ASRU7
2025 Interspeech 2025 URGENT Speech Enhancement Challenge
Kohei Saijo, Wangyou Zhang, Samuele Cornell, Robin Scheibler, Chenda Li, Zhaoheng Ni, Anurag Kumar 0003, Marvin Sach, Yihui Fu, Wei Wang 0010, Tim Fingscheidt, Shinji Watanabe 0001
INTERSPEECH4
2025 Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
Wangyou Zhang, Kohei Saijo, Samuele Cornell, Robin Scheibler, Chenda Li, Zhaoheng Ni, Anurag Kumar 0003, Marvin Sach, Wei Wang 0010, Yihui Fu, Shinji Watanabe 0001, Tim Fingscheidt, Yanmin Qian
INTERSPEECH4
2024 Universal Score-based Speech Enhancement with High Content Preservation
Robin Scheibler, Yusuke Fujita, Yuma Shirahata, Tatsuya Komatsu
INTERSPEECH1
2024 URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement
Wangyou Zhang, Robin Scheibler, Kohei Saijo, Samuele Cornell, Chenda Li, Zhaoheng Ni, Jan Pirklbauer, Marvin Sach, Shinji Watanabe 0001, Tim Fingscheidt, Yanmin Qian
INTERSPEECH2
2023 Neural Diarization with Non-Autoregressive Intermediate Attractors
abstract
End-to-end neural diarization (EEND) with encoder-decoder-based attractors (EDA) is a promising method to handle the whole speaker diarization problem simultaneously with a single neural network. While the EEND model can produce all frame-level speaker labels simultaneously, it disregards output label dependency. In this work, we propose a novel EEND model that introduces the label dependency between frames. The proposed method generates non-autoregressive intermediate attractors to produce speaker labels at the lower layers and conditions the subsequent layers with these labels. While the proposed model works in a non-autoregressive manner, the speaker labels are refined by referring to the whole sequence of intermediate labels. The experiments with the two-speaker CALLHOME dataset show that the intermediate labels with the proposed non-autoregressive intermediate attractors boost the diarization performance. The proposed method with the deeper net-work benefits more from the intermediate labels, resulting in better performance and training throughput than EEND-EDA.
Yusuke Fujita, Tatsuya Komatsu, Robin Scheibler, Yusuke Kida, Tetsuji Ogawa
ICASSP3
2023 Effectiveness of Inter- and Intra-Subarray Spatial Features for Acoustic Scene Classification
abstract
In this paper, we investigate the effectiveness of spatial features for acoustic scene classification (ASC) with distributed microphones. Assuming that multiple subarrays, each containing multiple micro-phones, are distributed and synchronized, we consider two types of generalized cross-correlation phase transform (GCC-PHAT) as spatial features: the intra- and inter-subarray GCC-PHATs. They are obtained from channels within the same subarray and between different subarrays, respectively. The log-Mel spectrogram as a spectral feature and the intra- or inter-subarray GCC-PHAT are processed in the neural network. The experimental results show that increasing the number of channels did not markedly improve the ASC performance when using the spectral features alone. However, using either of the GCC-PHATs as the spatial feature together with the spectral features successfully improved the ASC performance.
Takao Kawamura, Yuma Kinoshita, Nobutaka Ono, Robin Scheibler
ICASSP4
2023 Diffusion-Based Generative Speech Source Separation
abstract
We propose DiffSep, a new single channel source separation method based on score-matching of a stochastic differential equation (SDE). We craft a tailored continuous time diffusion-mixing process starting from the separated sources and converging to a Gaussian distribution centered on their mixture. This formulation lets us apply the machinery of score-based generative modelling. First, we train a neural network to approximate the score function of the marginal probabilities of the diffusion-mixing process. Then, we use it to solve the reverse time SDE that progressively separates the sources starting from their mixture. We propose a modified training strategy to handle model mismatch and source permutation ambiguity. Experiments on the WSJ0_2mix dataset demonstrate the potential of the method. Furthermore, the method is also suitable for speech enhancement and shows performance competitive with prior work on the VoiceBank-DEMAND dataset.
Robin Scheibler, Youna Ji, Soo-Whan Chung, Jaeuk Byun, Soyeon Choe, Min-Seok Choi
ICASSP1
2023 Multi-channel separation of dynamic speech and sound events
Takuya Fujimura, Robin Scheibler
INTERSPEECH2
2022 SDR - Medium Rare with Fast Computations
abstract
We revisit the widely used bss_eval metrics for source separation with an eye out for performance. We propose a fast algorithm fixing shortcomings of publicly available implementations. First, we show that the metrics are fully specified by the squared cosine of just two angles between estimate and reference subspaces. Second, large linear systems are involved. However, they are structured, and we apply a fast iterative method based on conjugate gradient descent. The complexity of this step is thus reduced by a factor quadratic in the distortion filter size used in bss_eval, usually 512. In experiments, we assess speed and numerical accuracy. Not only is the loss of accuracy due to the approximate solver acceptable for most applications, but the speed-up is up to two orders of magnitude in some, not so extreme, cases. We confirm that our implementation can train neural networks, and find that longer distortion filters may be beneficial.
Robin Scheibler
ICASSP1
2022 ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding
abstract
This paper presents recent progress on integrating speech separation and enhancement (SSE) into the ESPnet toolkit.Compared with the previous ESPnet-SE work, numerous features have been added, including recent state-of-the-art speech enhancement models with their respective training and evaluation recipes.Importantly, a new interface has been designed to flexibly combine speech enhancement front-ends with other tasks, including automatic speech recognition (ASR), speech translation (ST), and spoken language understanding (SLU).To showcase such integration, we performed experiments on carefully designed synthetic datasets for noisy-reverberant multichannel ST and SLU tasks, which can be used as benchmark corpora for future research.In addition to these new tasks, we also use CHiME-4 and WSJ0-2Mix to benchmark multiand single-channel SE approaches.Results show that the integration of SE front-ends with back-end tasks is a promising research direction even for tasks besides ASR, especially in the multi-channel scenario.The code is available online at https://github.com/ESPnet/ESPnet.The multichannel ST and SLU datasets, which are another contribution of this work, are released on HuggingFace.
Yen-Ju Lu, Xuankai Chang, Chenda Li, Wangyou Zhang, Samuele Cornell, Zhaoheng Ni, Yoshiki Masuyama, Brian Yan, Robin Scheibler, Zhongqiu Wang 0001, Yu Tsao 0001, Yanmin Qian, Shinji Watanabe 0001
INTERSPEECH9
2022 Independence-based Joint Dereverberation and Separation with Neural Source Model
abstract
We propose an independence-based joint dereverberation and separation method with a neural source model.We introduce a neural network in the framework of time-decorrelation iterative source steering, which is an extension of independent vector analysis to joint dereverberation and separation.The network is trained in an end-to-end manner with a permutation invariant loss on the time-domain separation output signals.Our proposed method can be applied in any situation with at least as many microphones as sources, regardless of their number.In experiments, we demonstrate that our method results in high performance in terms of both speech quality metrics and word error rate (WER), even for mixtures with a different number of speakers than training.Furthermore, the model, trained on synthetic mixtures, without any modifications, greatly reduces the WER on the recorded dataset LibriCSS.
Kohei Saijo, Robin Scheibler
INTERSPEECH2
2022 Spatial Loss for Unsupervised Multi-channel Source Separation
Kohei Saijo, Robin Scheibler
INTERSPEECH2
2022 End-to-End Multi-Speaker ASR with Independent Vector Analysis
abstract
We develop an end-to-end system for multi-channel, multi-speaker automatic speech recognition. We propose a frontend for joint source separation and dereverberation based on the independent vector analysis (IVA) paradigm. It uses the fast and stable iterative source steering algorithm together with a neural source model. Unlike conventional neural beamforming, the number of speakers can be dynamically changed during or after training. The parameters from the ASR module and the neural source model are optimized jointly from the ASR loss itself. We demonstrate competitive performance with previous systems using neural beamforming frontends with only one-ninth of the trainable parameter. First, we explore the trade-offs when using various number of channels for training and testing. Second, we demonstrate that the proposed IVA frontend performs well on noisy data, even when trained on clean mixtures only. Third, we demonstrate recognition of mixtures of three and four speakers with a model trained on mixtures of two only.
Robin Scheibler, Wangyou Zhang, Xuankai Chang, Shinji Watanabe 0001, Yanmin Qian
SLT1
2022 Computationally-Efficient Overdetermined Blind Source Separation Based on Iterative Source Steering
abstract
This paper describesa computationally-efficient optimization algorithm for the blind source separation (BSS) of overdetermined mixtures. In the determined case, a matrix-inversion-free iterative source steering (ISS) algorithm has been proposed for estimating a square demixing matrix as a computationally-efficient alternative to the popular iterative projection (IP) algorithm. The IP algorithm is based on source-wise (i.e., row-wise) updates of the demixing matrix, and lends itself naturally to an extension to overdetermined independent vector analysis (IVA) called OverIVA. In contrast, the ISS algorithm changes the whole demixing matrix at every update, making its extension to the overdetermined case non-trivial. In this paper, we propose a modified ISS algorithm for OverIVA fully exploiting the computational savings of ISS. We also derive an overdetermined extension of independent low-rank matrix analysis (OverILRMA) with the modified ISS algorithm. Experimental results showed that the proposed ISS-based OverIVA and OverILRMA were comparable or superior to the conventional IP-based counterparts in speech separation performance while achieving lower computational cost.
Yicheng Du, Robin Scheibler, Masahito Togami, Kazuyoshi Yoshii, Tatsuya Kawahara
IEEE Signal Process. Lett.2
2021 Joint Dereverberation and Separation With Iterative Source Steering
abstract
We propose a new algorithm for joint dereverberation and blind source separation (DR-BSS). Our work builds upon the IRLMA-T framework that applies a unified filter combining dereverberation and separation. One drawback of this framework is that it requires several matrix inversions, an operation inherently costly and with potential stability issues. We leverage the recently introduced iterative source steering (ISS) updates to propose two algorithms mitigating this issue. Albeit derived from first principles, the first algorithm turns out to be a natural combination of weighted prediction error (WPE) dereverberation and ISS-based BSS, applied alternatingly. In this case, we manage to reduce the number of matrix inversion to only one per iteration and source. The second algorithm updates the ILRMA-T matrix using only sequential ISS updates requiring no matrix inversion at all. Its implementation is straightforward and memory efficient. Numerical experiments demonstrate that both methods achieve the same final performance as ILRMA-T in terms of several relevant objective metrics. In the important case of two sources, the number of iterations required is also similar.
Taishi Nakashima, Robin Scheibler, Masahito Togami, Nobutaka Ono
ICASSP2
2021 Surrogate Source Model Learning for Determined Source Separation
abstract
We propose to learn surrogate functions of universal speech priors for determined blind speech separation. Deep speech priors are highly desirable due to their superior modelling power, but are not compatible with state-of-the-art independent vector analysis based on majorization-minimization (AuxIVA), since deriving the required surrogate function is not easy, nor always possible. Instead, we do away with exact majorization and directly approximate the surrogate. Taking advantage of iterative source steering (ISS) updates, we back propagate the permutation invariant separation loss through multiple iterations of AuxIVA. ISS lends itself well to this task due to its lower complexity and lack of matrix inversion. Experiments show large improvements in terms of scale invariant signal-to-distortion (SDR) ratio and word error rate compared to baseline methods. Training is done on two speakers mixtures and we experiment with two losses, SDR and coherence. We find that the learnt approximate surrogate generalizes well on mixtures of three and four speakers without any modification. We also demonstrate generalization to a different variation of the AuxIVA update equations. The SDR loss leads to fastest convergence in iterations, while coherence leads to the lowest word error rate (WER). We obtain as much as 36 % reduction in WER.
Robin Scheibler, Masahito Togami
ICASSP1
2021 Refinement of Direction of Arrival Estimators by Majorization-Minimization Optimization on the Array Manifold
abstract
We propose a generalized formulation of direction of arrival estimation that includes many existing methods such as steered response power, subspace, coherent and incoherent, as well as speech sparsity-based methods. Unlike most conventional methods that rely exclusively on grid search, we introduce a continuous optimization algorithm to refine DOA estimates beyond the resolution of the initial grid. The algorithm is derived from the majorization-minimization (MM) technique. We derive two surrogate functions, one quadratic and one linear. Both lead to efficient iterative algorithms that do not require hyperparameters, such as step size, and ensure that the DOA estimates never leave the array manifold, without the need for a projection step. In numerical experiments, we show that the accuracy after a few iterations of the MM algorithm nearly removes dependency on the resolution of the initial grid used. We find that the quadratic surrogate function leads to very fast convergence, but the simplicity of the linear algorithm is very attractive, and the performance gap small.
Robin Scheibler, Masahito Togami
ICASSP1
2021 Sound Source Localization with Majorization Minimization
Masahito Togami, Robin Scheibler
Interspeech2
2020 Fast and Stable Blind Source Separation with Rank-1 Updates
abstract
We propose a new algorithm for the blind source separation of acoustic sources. This algorithm is an alternative to the popular auxiliary function based independent vector analysis using iterative projection (AuxIVA-IP). It optimizes the same cost function, but instead of alternate updates of the rows of the demixing matrix, we propose a sequence of rank-1 updates. Remarkably, and unlike the previous method, the resulting updates do not require matrix inversion. Moreover, their computational complexity is quadratic in the number of microphones, rather than cubic in AuxIVA-IP. In addition, we show that the new method can be derived as alternate updates of the steering vectors of sources. Accordingly, we name the method iterative source steering (AuxIVA-ISS). Finally, we confirm in simulated experiments that the proposed algorithm separates sources just as well as AuxIVA-IP, at a lower computational cost.
Robin Scheibler, Nobutaka Ono
ICASSP1
2020 Fast Independent Vector Extraction by Iterative SINR Maximization
abstract
We propose fast independent vector extraction (FIVE), a new algorithm that blindly extracts a single non-Gaussian source from a Gaussian background. The algorithm iteratively computes beam-forming weights maximizing the signal-to-interference-and-noise ratio for an approximate noise covariance matrix. We demonstrate that this procedure minimizes the negative log-likelihood of the input data according to a well-defined probabilistic model. The minimization is carried out via the auxiliary function technique whereas, unlike related methods, the auxiliary function is globally minimized at every iteration. Numerical experiments are carried out to assess the performance of FIVE. We find that it is vastly superior to competing methods in terms of convergence speed, and has high potential for real-time applications.
Robin Scheibler, Nobutaka Ono
ICASSP1
2020 Generalized Minimal Distortion Principle for Blind Source Separation
abstract
We revisit the source image estimation problem from blind source separation (BSS). We generalize the traditional minimum distortion principle to maximum likelihood estimation with a model for the residual spectrograms. Because residual spectrograms typically contain other sources, we propose to use a mixed-norm model that lets us finely tune sparsity in time and frequency. We propose to carry out the minimization of the mixed-norm via majorization-minimization optimization, leading to an iteratively reweighted least-squares algorithm. The algorithm balances well efficiency and ease of implementation. We assess the performance of the proposed method as applied to two well-known determined BSS and one joint BSS-dereverberation algorithms. We find out that it is possible to tune the parameters to improve separation by up to 2 dB, with no increase in distortion, and at little computational cost. The method thus provides a cheap and easy way to boost the performance of blind source separation.
Robin Scheibler
INTERSPEECH1
2020 Sparseness-Aware DOA Estimation with Majorization Minimization
Masahito Togami, Robin Scheibler
INTERSPEECH2
2019 Multi-modal Blind Source Separation with Microphones and Blinkies
abstract
We propose a blind source separation algorithm that jointly exploits measurements by a conventional microphone array and an ad hoc array of low-rate sound power sensors called blinkies. While providing less information than microphones, blinkies circumvent some difficulties of microphone arrays in terms of manufacturing, synchronization, and deployment. The algorithm is derived from a joint probabilistic model of the microphone and sound power measurements. We assume the separated sources to follow a time-varying spherical Gaussian distribution, and the non-negative power measurement space-time matrix to have a low-rank structure. We show that alternating updates similar to those of independent vector analysis and Itakura-Saito non-negative matrix factorization decrease the negative log-likelihood of the joint distribution. The proposed algorithm is validated via numerical experiments. Its median separation performance is found to be up to 8 dB more than that of independent vector analysis, with significantly reduced variability.
Robin Scheibler, Nobutaka Ono
ICASSP1
2019 Blink-former: Light-aided beamforming for multiple targets enhancement
abstract
We propose a multimodal framework to enhance multiple target sound sources using a conventional microphone array, a video camera, and sound power sensors, called Blinkies, that we have recently developed. Each Blinky consists of a microphone, LEDs, a microcontroller, and a battery. One of the LEDs intensity is varied according to sound power, that is, the Blinky works as a sound-to-light conversion sensor. They are easy to distribute over a large area, and thus, the sound power information therein can be harvested by capturing the LED signals with a video camera. Although these signals are a mixture of contributions from multiple sources, we demonstrate that they can be separated into individual source activities by non-negative matrix factorization. The obtained activities are further utilized to design maximum signal-to-interference-and-noise ratio beamformers enhancing the source signals. We conduct numerical simulations and real experiments to evaluate the performance of this method in diffuse noise environment. The experimental results show that the proposed scheme using Blinkies is superior to competing algorithms, especially at low signal-to-noise ratio.
Daiki Horiike, Robin Scheibler, Yukoh Wakabayashi, Nobutaka Ono
MMSP2
2018 Combining Range and Direction for Improved Localization
abstract
Self-localization of nodes in a sensor network is typically achieved using either range or direction measurements; in this paper, we show that a constructive combination of both improves the estimation. We propose two localization algorithms that make use of the differences between the sensors' coordinates, or edge vectors; these can be calculated from measured distances and angles. Our first method improves the existing edge-multidimensional scaling algorithm (E- MDS) by introducing additional constraints that enforce geometric consistency between the edge vectors. On the other hand, our second method decomposes the edge vectors onto 1-dimensional spaces and introduces the concept of coordinate difference matrices (CDMs) to independently regularize each projection. This solution is optimal when Gaussian noise is added to the edge vectors. We demonstrate in numerical simulations that both algorithms outperform state-of-the-art solutions.
Gilles Baechler, Frederike Diimbgen, Golnoosh Elhami, Miranda Krekovic, Robin Scheibler, Adam Scholefield, Martin Vetterli
ICASSP5
2018 Pyroomacoustics: A Python Package for Audio Room Simulation and Array Processing Algorithms
abstract
We present pyroomacoustics, a software package aimed at the rapid development and testing of audio array processing algorithms. The content of the package can be divided into three main components: an intuitive Python object-oriented interface to quickly construct different simulation scenarios involving multiple sound sources and microphones in 2D and 3D rooms; a fast C implementation of the image source model for general polyhedral rooms to efficiently generate room impulse responses and simulate the propagation between sources and receivers; and finally, reference implementations of popular algorithms for beamforming, direction finding, and adaptive filtering. Together, they form a package with the potential to speed up the time to market of new algorithms by significantly reducing the implementation overhead in the performance evaluation step.
Robin Scheibler, Eric Bezzam, Ivan Dokmanic
ICASSP1
2018 Separake: Source Separation with a Little Help from Echoes
abstract
It is commonly believed that multipath hurts various audio processing algorithms. At odds with this belief, we show that multipath in fact helps sound source separation, even with very simple propagation models. Unlike most existing methods, we neither ignore the room impulse responses, nor we attempt to estimate them fully. We rather assume to know the positions of a few virtual microphones generated by echoes and we show how this gives us enough spatial diversity to get a performance boost over the anechoic case. We show improvements for two standard algorithms-one that uses only magnitudes of the transfer functions, and one that also uses the phases. Concretely, we show that multi-channel non-negative matrix factorization aided with a small number of echoes beats the vanilla variant of the same algorithm, and that with magnitude information only, echoes enable separation where it was previously impossible.
Robin Scheibler, Diego Di Carlo, Antoine Deleforge, Ivan Dokmanic
ICASSP1
2017 Hardware and software for reproducible research in audio array signal processing
abstract
In our demo, we present two hardware platforms for prototyping audio array signal processing. Pyramic is a 48-channel microphone array fitted on an FPGA and Compact Six is a portable microphone array with six microphones, closer to the technical constraints of consumer electronics. A browser based interface was developed that allows the user to interact with the audio stream from the arrays in real time. The software component of this demo is a Python module with implementations of basic audio signal processing blocks and popular techniques like STFT, beamforming, and DoA. Both the hardware design files and the software are open source and freely shared. As part of a collaboration with IBM Research, their beamforming and imaging technologies will also be portrayed. The hardware will be demonstrated through an installation processing the microphone signals into light patterns on a circular LED array. The demo will be interactive and let visitors play with different algorithms for DoA (SRP, FRIDA [1], Bluebild) and beamforming (MVDR, Flexibeam [2]). The availability of an open platform with reference implementations encourages reproducible research and minimizes setup-time when testing and benchmarking new audio array signal processing algorithms. It can also serve as a useful educational tool, providing a means to work with real-life signals.
Eric Bezzam, Robin Scheibler, Juan Azcarreta, Hanjie Pan, Matthieu Simeoni, Rene Beuchat, Paul Hurley, Basile Bruneau, Corentin Ferry, Sepand Kashani
ICASSP2
2017 FRIDA: FRI-based DOA estimation for arbitrary array layouts
abstract
In this paper we present FRIDA-an algorithm for estimating directions of arrival of multiple wideband sound sources. FRIDA combines multi-band information coherently and achieves state-of-the-art resolution at extremely low signal-to-noise ratios. It works for arbitrary array layouts, but unlike the various steered response power and subspace methods, it does not require a grid search. FRIDA leverages recent advances in sampling signals with a finite rate of innovation. It is based on the insight that for any array layout, the entries of the spatial covariance matrix can be linearly transformed into a uniformly sampled sum of sinusoids.
Hanjie Pan, Robin Scheibler, Eric Bezzam, Ivan Dokmanic, Martin Vetterli
ICASSP2
2016 The recursive hessian sketch for adaptive filtering
abstract
We introduce in this paper the recursive Hessian sketch, a new adaptive filtering algorithm based on sketching the same exponentially weighted least squares problem solved by the recursive least squares algorithm. The algorithm maintains a number of sketches of the inverse autocorrelation matrix and recursively updates them at random intervals. These are in turn used to update the unknown filter estimate. The complexity of the proposed algorithm compares favorably to that of recursive least squares. The convergence properties of this algorithm are studied through extensive numerical experiments. With an appropriate choice or parameters, its convergence speed falls between that of least mean squares and recursive least squares adaptive filters, with less computations than the latter.
Robin Scheibler, Martin Vetterli
ICASSP1
2015 Raking echoes in the time domain
abstract
The geometry of room acoustics is such that the reverberant signal can be seen as the same waveform emitted from multiple locations. In analogy with the rake receiver from wireless communications, we propose several beamforming strategies that exploit, rather than suppress, this additional spatio-temporal diversity. Unlike earlier work in the frequency domain, time domain designs allow to shape the impulse response of the beamformer. In particular, we can control perceptually relevant parameters, such as the amount of early echoes or the length of the beamformer response. Relying on the knowledge of the image sources positions, we derive different optimal beamformers. Leveraging perceptual cues, we show how to improve interference and noise reduction without degrading the perceptual quality. The designs are validated through simulation. Using early echoes is shown to strictly improve the signal to interference and noise ratio. Code and speech samples are available online at http:// lcav.epfl.ch/Robin_Scheibler.
Robin Scheibler, Ivan Dokmanic, Martin Vetterli
ICASSP1
2015 A Fast Hadamard Transform for Signals With Sublinear Sparsity in the Transform Domain
abstract
In this paper, we design a new iterative low-complexity algorithm for computing the Walsh-Hadamard transform (WHT) of an N dimensional signal with a K-sparse WHT. We suppose that N is a power of two and K = O(Nα), scales sublinearly in N for some α ∈ (0, 1). Assuming a random support model for the nonzero transform-domain components, our algorithm reconstructs the WHT of the signal with a sample complexity O(K log2(N/K)) and a computational complexity O(K log2(K) log2(N/K)). Moreover, the algorithm succeeds with a high probability approaching 1 for large dimension N. Our approach is mainly based on the subsampling (aliasing) property of the WHT, where by a carefully designed subsampling of the time-domain signal, a suitable aliasing pattern is induced in the transform domain. We treat the resulting aliasing patterns as parity-check constraints and represent them by a bipartite graph. We analyze the properties of the resulting bipartite graphs and borrow ideas from codes defined over sparse bipartite graphs to formulate the recovery of the nonzero spectral values as a peeling decoding algorithm for a specific sparse-graph code transmitted over a binary erasure channel. This enables us to use tools from coding theory (belief-propagation analysis) to characterize the asymptotic performance of our algorithm in the very sparse (α ∈ (0, 1/3]) and the less sparse (α ∈ (1/3, 1)) regime. Comprehensive simulation results are provided to assess the empirical performance of the proposed algorithm.
Robin Scheibler, Saeid Haghighatshoar, Martin Vetterli
IEEE Trans. Inf. Theory1
2013 The Fukushima inverse problem
abstract
Knowing what amount of radioactive material was released from Fukushima in March 2011 is crucial to understand the scope of the consequences. Moreover, it could be used in forward simulations to obtain accurate maps of deposition. But these data are often not publicly available, or are of questionable quality. We propose to estimate the emission waveforms by solving an inverse problem. Previous approaches rely on a detailed expert guess of how the releases appeared, and they produce a solution strongly biased by this guess. If we plant a nonexistent peak in the guess, the solution also exhibits a nonexistent peak. We propose a method based on sparse regularization that solves the Fukushima inverse problem blindly. Together with the atmospheric dispersion models and worldwide radioactivity measurements our method correctly reconstructs the times of major events during the accident, and gives plausible estimates of the released quantities of Xenon.
Marta Martinez-Camara, Ivan Dokmanic, Juri Ranieri, Robin Scheibler, Martin Vetterli, Andreas Stohl
ICASSP4