VLDB 2026 Research / reviewers in the wild / expert
Nicki Holighaus
dblp:02/10577
· DBLP profile ↗
9ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0003-3837-2865ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust Multicomponent Tracking of Ultrasonic VocalizationsabstractUltrasonic vocalizations (USV) convey information about individual identity and arousal status in mice. We propose to track USV as ridges in the time–frequency domain via a variant of time– frequency reassignment (TFR). The key idea is to perform TFR with empirical Wiener shrinkage and multitapering to improve robustness to noise. Furthermore, we perform TFR over both the short-term Fourier transform and the constant-Q transform so as to detect both the fundamental frequency and its harmonic partial (if any). Experimental results show that our approach effectively estimates multicomponent ridges with high precision and low frequency deviation. Reyhaneh Abbasi, Nicki Holighaus, Vincent Lostanlen, Dustin J. Penn, Sarah M. Zala |
ICASSP | 2 |
| 2023 | Grid-Based Decimation for Wavelet Transforms With Stably Invertible ImplementationabstractThe constant center frequency to bandwidth ratio (Q-factor) of wavelet transforms provides a very natural representation for audio data. However, invertible wavelet transforms have either required non-uniform decimation—leading to irregular data structures that are cumbersome to work with—or require excessively high oversampling with unacceptable computational overhead. Here, we present a novel decimation strategy for wavelet transforms that leads to stable representations with oversampling rates close to one and uniform decimation. Specifically, we show that finite implementations of the resulting representation are energy-preserving in the sense of frame theory. The obtained wavelet coefficients can be stored in a time-frequency matrix with a natural interpretation of columns as time frames and rows as frequency channels. This matrix structure immediately grants access to a large number of algorithms that are successfully used in time-frequency audio processing, but could not previously be used jointly with wavelet transforms. We demonstrate the application of our method in processing based on nonnegative matrix factorization, in onset detection, and in phaseless reconstruction. Nicki Holighaus, Günther Koliander, Clara Hollomey, Friedrich Pillichshammer |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Fast Matching Pursuit with Multi-Gabor DictionariesabstractFinding the best K -sparse approximation of a signal in a redundant dictionary is an NP-hard problem. Suboptimal greedy matching pursuit algorithms are generally used for this task. In this work, we present an acceleration technique and an implementation of the matching pursuit algorithm acting on a multi-Gabor dictionary, i.e., a concatenation of several Gabor-type time-frequency dictionaries, each of which consists of translations and modulations of a possibly different window and time and frequency shift parameters. The technique is based on pre-computing and thresholding inner products between atoms and on updating the residual directly in the coefficient domain, i.e., without the round-trip to the signal domain. Since the proposed acceleration technique involves an approximate update step, we provide theoretical and experimental results illustrating the convergence of the resulting algorithm. The implementation is written in C (compatible with C99 and C++11), and we also provide Matlab and GNU Octave interfaces. For some settings, the implementation is up to 70 times faster than the standard Matching Pursuit Toolkit. Zdenek Prusa, Nicki Holighaus |
ACM Trans. Math. Softw. | 2 |
| 2019 | Adversarial Generation of Time-Frequency Features with application in audio synthesisabstractTime-frequency (TF) representations provide powerful and intuitive features for the analysis of time series such as audio. But still, generative modeling of audio in the TF domain is a subtle matter. Consequently, neural audio synthesis widely relies on directly modeling the waveform and previous attempts at unconditionally synthesizing audio from neurally generated invertible TF features still struggle to produce audio at satisfying quality. In this article, focusing on the short-time Fourier transform, we discuss the challenges that arise in audio synthesis based on generated invertible TF features and how to overcome them. We demonstrate the potential of deliberate generative TF modeling by training a generative adversarial network (GAN) on short-time Fourier features. We show that by applying our guidelines, our TF-based network was able to outperform a state-of-the-art GAN generating waveforms directly, despite the similar architecture in the two networks. Andrés Marafioti, Nathanaël Perraudin, Nicki Holighaus, Piotr Majdak |
ICML | 3 |
| 2019 | A Context Encoder For Audio InpaintingabstractIn this article, we study the ability of deep neural networks (DNNs) to restore missing audio content based on its context, i.e., inpaint audio gaps. We focus on a condition which has not received much attention yet: gaps in the range of tens of milliseconds. We propose a DNN structure that is provided with the signal surrounding the gap in the form of time-frequency (TF) coefficients. Two DNNs with either complex-valued TF coefficient output or magnitude TF coefficient output were studied by separately training them on inpainting two types of audio signals (music and musical instruments) having 64-ms long gaps. The magnitude DNN outperformed the complex-valued DNN in terms of signal-to-noise ratios and objective difference grades. Although, for instruments, a reference inpainting obtained through linear predictive coding performed better in both metrics, it performed worse than the magnitude DNN for music. This demonstrates the potential of the magnitude DNN, in particular for inpainting signals that are more complex than single instrument sounds. Andrés Marafioti, Nathanaël Perraudin, Nicki Holighaus, Piotr Majdak |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Inpainting of Long Audio Segments With Similarity GraphsabstractPre-trained models for:\n\nAdversarial Generation of Time-Frequency Features with application in audio synthesis\nA Marafioti, N Holighaus, N Perraudin, P Majdak - arXiv preprint arXiv:1902.04072, 2019\n\n Nathanaël Perraudin, Nicki Holighaus, Piotr Majdak |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Reassignment and synchrosqueezing for general time-frequency filter banks, subsampling and processing
Nicki Holighaus, Zdenek Prusa, Peter L. Søndergaard |
Signal Process. | 1 |
| 2013 | The ERBlet transform: An auditory-based time-frequency representation with perfect reconstructionabstractThis paper describes a method for obtaining a perceptually motivated and perfectly invertible time-frequency representation of a sound signal. Based on frame theory and the recent non-stationary Gabor transform, a linear representation with resolution evolving across frequency is formulated and implemented as a non-uniform filterbank. To match the human auditory time-frequency resolution, the transform uses Gaussian windows equidistantly spaced on the psychoacoustic “ERB” frequency scale. Additionally, the transform features adaptable resolution and redundancy. Simulations showed that perfect reconstruction can be achieved using fast iterative methods and preconditioning even using one filter per ERB and a very low redundancy (1.08). Comparison with a linear gammatone filterbank showed that the ERBlet approximates well the auditory time-frequency resolution. Thibaud Necciari, Nicki Holighaus, Peter L. Søndergaard |
ICASSP | 3 |
| 2013 | A Framework for Invertible, Real-Time Constant-Q TransformsabstractAudio signal processing frequently requires time-frequency representations and in many applications, a non-linear spacing of frequency bands is preferable. This paper introduces a framework for efficient implementation of invertible signal transforms allowing for non-uniform frequency resolution. Non-uniformity in frequency is realized by applying nonstationary Gabor frames with adaptivity in the frequency domain. The realization of a perfectly invertible constant-Q transform is described in detail. To achieve real-time processing, independent of signal length, slice-wise processing of the full input signal is proposed and referred to as sliCQ transform. By applying frame theory and FFT-based processing, the presented approach overcomes computational inefficiency and lack of invertibility of classical constant-Q transform implementations. Numerical simulations evaluate the efficiency of the proposed algorithm and the method's applicability is illustrated by experiments on real-life audio signals . Nicki Holighaus, Monika Dörfler, Gino Angelo M. Velasco, Thomas Grill |
IEEE Trans. Speech Audio Process. | 1 |