Kevin Wilkinghoff

dblp:207/9559 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0003-4200-9129ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 No Class Left Behind: A Closer Look at Class Balancing for Audio Tagging
abstract
Large-scale audio tagging datasets like AudioSet usually suffer from severe class imbalance comprising many audio examples for common sound classes but only few examples of rare sound classes. The latter, however, may yet be equally or even more important to recognize. Therefore, it is common practice to sample examples from rare classes more frequently during training. At the same time, the effects of such balancing on a model’s training and tagging performance are still little understood. In this work, we investigate how it affects training convergence and tagging performance. We consider varying degrees of balancing and investigate whether classes converge simultaneously or if there is a benefit from selecting different balancing rates for each class. Furthermore, we investigate data efficient oversampling, which keeps audio files from rare classes in memory, and repeats them in close succession over multiple batches, minimizing data loading from disk. Finally, we show that for AudioSet, the optimal amount of class balancing is different when fine-tuning a model pre-trained via self-supervised learning, versus training a supervised model from scratch.
Janek Ebbers, François G. Germain, Kevin Wilkinghoff, Gordon Wichern, Jonathan Le Roux
ICASSP3
2025 Keeping the Balance: Anomaly Score Calculation for Domain Generalization
abstract
Emitted sounds may drastically change when using different microphones, when properties of the sound sources change, or when recording in different acoustic environments. Ideally, anomalous sound detection (ASD) systems should be able to generalize well to unseen target domains by only providing a few target domain samples to define how normal data samples sound like, without needing to re-train or modify the system. In contrast with the source domain, for which many normal training samples are available, accurately estimating the underlying distribution of normal data after a domain shift based on very few samples is challenging. This usually leads to a mismatch between the corresponding anomaly scores of source and target domains and significantly reduces performance. In this work, we propose a framework for re-scaling anomaly scores based on the ratio between the cosine distance of a test sample to a normal reference sample and the distances to this sample’s next-closest neighbors in the reference set. In experimental evaluations, it is shown that the re-scaled anomaly scores reduce the domain mismatch for multiple domains. As a result, we obtain new state-of-the-art performances on the DCASE2020 and DCASE2023 ASD datasets.
Kevin Wilkinghoff, Haici Yang, Janek Ebbers, François G. Germain, Gordon Wichern, Jonathan Le Roux
ICASSP1
2024 Self-Supervised Learning for Anomalous Sound Detection
abstract
State-of-the-art anomalous sound detection (ASD) systems are often trained by using an auxiliary classification task to learn an embedding space. Doing so enables the system to learn embeddings that are robust to noise and are ignoring non-target sound events but requires manually annotated meta information to be used as class labels. However, the less difficult the classification task becomes, the less informative are the embeddings and the worse is the resulting ASD performance. A solution to this problem is to utilize selfsupervised learning (SSL). In this work, feature exchange (FeatEx), a simple yet effective SSL approach for ASD, is proposed. In addition, FeatEx is compared to and combined with existing SSL approaches. As the main result, a new state-of-the-art performance for the DCASE2023 ASD dataset is obtained that outperforms all other published results on this dataset by a large margin.
Kevin Wilkinghoff
ICASSP1
2024 TACos: Learning Temporally Structured Embeddings for Few-Shot Keyword Spotting with Dynamic Time Warping
abstract
To segment a signal into blocks to be analyzed, few-shot keyword spotting (KWS) systems often utilize a sliding window of fixed size. Because of the varying lengths of different keywords or their spoken instances, choosing the right window size is a problem: A window should be long enough to contain all necessary information needed to recognize a keyword but a longer window may contain irrelevant information such as multiple words or noise and thus makes it difficult to reliably detect on- and offsets of keywords. We propose TACos, a novel angular margin loss for deriving two-dimensional embeddings that retain temporal properties of the underlying speech signal. In experiments conducted on KWS-DailyTalk, a few-shot KWS dataset presented in this work, using these embeddings as templates for dynamic time warping is shown to outperform using other representations or a sliding window and that using time-reversed segments of the keywords during training improves the performance.
Kevin Wilkinghoff, Alessia Cornaggia-Urrigshardt
ICASSP1
2024 F1-EV score: Measuring The Likelihood of Estimating a Good Decision Threshold for Semi-Supervised Anomaly Detection
abstract
Anomalous sound detection (ASD) systems are usually compared by using threshold-independent performance measures such as AUCROC. However, for practical applications a decision threshold is needed to decide whether a given test sample is normal or anomalous. Estimating such a threshold is highly non-trivial in a semi-supervised setting where only normal training samples are available. In this work, F1-EV a novel threshold-independent performance measure for ASD systems that also includes the likelihood of estimating a good decision threshold is proposed and motivated using specific toy examples. In experimental evaluations, multiple performance measures are evaluated for all systems submitted to the ASD task of the DCASE Challenge 2023. It is shown that F1-EV is strongly correlated with AUC-ROC while having a significantly stronger correlation with the F1-score obtained with estimated and optimal decision thresholds than AUC-ROC.
Kevin Wilkinghoff, Keisuke Imoto
ICASSP1
2024 Why Do Angular Margin Losses Work Well for Semi-Supervised Anomalous Sound Detection?
abstract
State-of-the-art anomalous sound detection systems often utilize angular margin losses to learn suitable representations of acoustic data using an auxiliary task, which usually is a supervised or self-supervised classification task. The underlying idea is that, in order to solve this auxiliary task, specific information about normal data needs to be captured in the learned representations and that this information is also sufficient to differentiate between normal and anomalous samples. Especially in noisy conditions, discriminative models based on angular margin losses tend to significantly outperform systems based on generative or one-class models. The goal of this work is to investigate why using angular margin losses with auxiliary tasks works well for detecting anomalous sounds. To this end, it is shown, both theoretically and experimentally, that minimizing angular margin losses also minimizes compactness loss while inherently preventing learning trivial solutions. Furthermore, multiple experiments are conducted to show that using a related classification task as an auxiliary task teaches the model to learn representations suitable for detecting anomalous sounds in noisy conditions. Among these experiments are performance evaluations, visualizing the embedding space with t-SNE and visualizing the input representations with respect to the anomaly score using randomized input sampling for explanation.
Kevin Wilkinghoff, Frank Kurth
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Design Choices for Learning Embeddings from Auxiliary Tasks for Domain Generalization in Anomalous Sound Detection
abstract
Emitted machine sounds can change drastically due to a change in settings of machines or varying noise conditions resulting in false alarms when monitoring machine conditions with a trained anomalous sound detection (ASD) system. In this work, a conceptually simple state-of-the-art ASD system based on embeddings learned through auxiliary tasks generalizing to multiple data domains is presented. In experiments conducted on the DCASE 2022 ASD dataset, particular design choices such as preventing trivial projections, combining multiple input representations and choosing a suitable back-end are shown to significantly improve the ASD performance.
Kevin Wilkinghoff
ICASSP1
2021 Sub-Cluster AdaCos: Learning Representations for Anomalous Sound Detection
abstract
When training a model for anomalous sound detection, one usually needs to estimate the underlying distribution of the normal data. By doing so, anomalous data has a lower probability in view of this distribution than normal data and thus can easily be detected. However, audio data is very high-dimensional making it difficult to have a good estimate of the true distribution. To have more accurate estimates, the dimension of the data can be reduced first. One way to do this is to train discriminative neural networks for extracting lower-dimensional representations of the data. Particularly, neural networks trained with angular margin losses as AdaCos have been shown to perform well for this task. In this work, a modified AdaCos loss called sub-cluster AdaCos specifically designed for detecting anomalous data is presented. In multiple experiments conducted on the DCASE 2020 dataset for “Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring”, these design choices are empirically justified. As a result, a conceptually simple system for anomalous sound detection is presented that significantly outperforms all other published systems on this dataset.
Kevin Wilkinghoff
IJCNN1
2018 Robust Detection of Jittered Multiply Repeating Audio Events Using Iterated Time-Warped ACF
abstract
This paper proposes a novel approach for robustly detecting multiply repeating audio events in monitoring recordings. We consider the practically important case that the sequence of inter onset intervals between subsequent events is not constant but differs by some jitter. In such cases classical approaches based on autocorrelation (ACF) are of limited use. To overcome this problem we propose to use ACF together with a variant of dynamic time warping. Combining both techniques in an iterative algorithm, we obtain a method for significantly improved detection of jittered multiply repeating events. In this paper we describe the new iterated time-warped ACF algorithm and evaluate its performance on the bioacoustic application of detecting repeating bird calls in monitoring recordings.
Frank Kurth, Kevin Wilkinghoff
ICASSP2