Nils Peters

dblp:45/9005 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Why disentanglement-based speaker anonymization systems fail at preserving emotions?
abstract
Disentanglement-based speaker anonymization involves decomposing speech into a semantically meaningful representation, altering the speaker embedding, and resynthesizing a waveform using a neural vocoder. State-of-the-art systems of this kind are known to remove emotion information. Possible reasons include mode collapse in GAN-based vocoders, unintended modeling and modification of emotions through speaker embeddings, or excessive sanitization of the intermediate representation (IR). In this paper, we conduct a comprehensive evaluation of a state-of-the-art speaker anonymization system to understand the underlying causes. We conclude that the main reason is the lack of emotion-related information in the IR. The speaker embeddings also have a high impact, if they are learned in a generative context. The vocoder’s out-of-distribution performance has a smaller impact. Additionally, we discovered that synthesis artifacts increase spectral kurtosis, biasing emotion recognition evaluation towards classifying utterances as angry. Therefore, we conclude that reporting unweighted average recall alone for emotion recognition performance is suboptimal.
Ünal Ege Gaznepoglu, Nils Peters
ICASSP2
2025 Acoustic Position Estimation of a Silent Listener
abstract
The localization of sounding objects has been extensively studied. In certain situations, however, an object of interest may not emit sound while another sound source is active. For instance, this occurs when a silent human listens to music and needs to be localized, e.g., to optimize the playback for the listener’s position. In this study, we present a method for localizing silent listeners: By continuously estimating room impulse responses using a loudspeaker emitting a known signal and a microphone array, we can detect subtle variations in the room impulse response over time, which are induced by the listener’s motion. Our results, evaluated in two different rooms, indicate that six room impulse responses, estimated using two successive 20 Hz-24 kHz sine sweeps of 1 s duration and each captured at three spatial points, are sufficient to localize a sitting listener within 35 cm accuracy 82 % and a standing listener 93 % of the time.
Jeremy Lawrence, Cagdas Tuna, Andreas Walther 0001, Nils Peters
ICASSP4
2025 Concentrating Harder for Faster Audio Transformer
abstract
Attention-based models have become tremendously successful in the last couple of years for tasks such as Acoustic Scene Classification, Event Classification, or Speaker Identification. They operate on a set of tokens extracted from audio features and scale quadratically in sequence length due to their pairwise operations in multi-head attention.In this paper, we propose to parameterize a Dirichlet prior with a Transformer model, to jointly estimate label and token probabilities. A token bottleneck, provided by a Categorical-Dirichlet pair, forces the model to concentrate on a subset of tokens. This allows for improved interpretability and higher audio throughput during inference. We compare two different methods for token sampling – full knowledge of all tokens and partial token sampling for reduced complexity.We evaluate and interpret typical audio datasets such as the Environmental Sound Classification (ESC-50) dataset, the TAU Urban Acoustic Scenes 2020 (TAU20) dataset, and the Speech Command version 2 (SCv2) dataset. The results show that the token budget can be reduced without significant performance loss, especially for acoustic scene classification. We show that for the ESC-50 and SCv2 datasets, the token relevance can be well approximated with partial token view. Finally, we show that a significant increase in throughput can be achieved with our proposed methods.
Lorenz P. Schmidt, Nils Peters
ICASSP2
2025 You Are What You Say: Exploiting Linguistic Content for VoicePrivacy Attacks
abstract
4238
Ünal Ege Gaznepoglu, Anna Leschanowsky, Ahmad Aloradi, Prachi Singh, Daniel Tenbrinck, Emanuël A. P. Habets, Nils Peters
INTERSPEECH7
2024 Voice Anonymization for All-Bias Evaluation of the Voice Privacy Challenge Baseline Systems
abstract
In an age of voice-enabled technology, voice anonymization offers a solution to protect people’s privacy, provided these systems work equally well across subgroups. This study investigates bias in voice anonymization systems within the context of the Voice Privacy Challenge. We curate a novel benchmark dataset to assess performance disparities among speaker subgroups based on sex and dialect. We analyze the impact of three anonymization systems and attack models on speaker subgroup bias and reveal significant performance variations. Notably, subgroup bias intensifies with advanced attacker capabilities, emphasizing the challenge of achieving equal performance across all subgroups. Our study highlights the need for inclusive benchmark datasets and comprehensive evaluation strategies that address subgroup bias in voice anonymization.
Anna Leschanowsky, Ünal Ege Gaznepoglu, Nils Peters
ICASSP3
2024 A Denoising Diffusion Probabilistic Model for Metal Artifact Reduction in CT
abstract
The presence of metal objects leads to corrupted CT projection measurements, resulting in metal artifacts in the reconstructed CT images. AI promises to offer improved solutions to estimate missing sinogram data for metal artifact reduction (MAR), as previously shown with convolutional neural networks (CNNs) and generative adversarial networks (GANs). Recently, denoising diffusion probabilistic models (DDPM) have shown great promise in image generation tasks, potentially outperforming GANs. In this study, a DDPM-based approach is proposed for inpainting of missing sinogram data for improved MAR. The proposed model is unconditionally trained, free from information on metal objects, which can potentially enhance its generalization capabilities across different types of metal implants compared to conditionally trained approaches. The performance of the proposed technique was evaluated and compared to the state-of-the-art normalized MAR (NMAR) approach as well as to CNN-based and GAN-based MAR approaches. The DDPM-based approach provided significantly higher SSIM and PSNR, as compared to NMAR (SSIM: p [Formula: see text]; PSNR: p [Formula: see text]), the CNN (SSIM: p [Formula: see text]; PSNR: p [Formula: see text]) and the GAN (SSIM: p [Formula: see text]; PSNR: p <0.05) methods. The DDPM-MAR technique was further evaluated based on clinically relevant image quality metrics on clinical CT images with virtually introduced metal objects and metal artifacts, demonstrating superior quality relative to the other three models. In general, the AI-based techniques showed improved MAR performance compared to the non-AI-based NMAR approach. The proposed methodology shows promise in enhancing the effectiveness of MAR, and therefore improving the diagnostic accuracy of CT.
Grigorios M. Karageorgos, Jiayong Zhang, Nils Peters, Wenjun Xia, Chuang Niu, Harald Paganetti, Ge Wang 0001, Bruno De Man
IEEE Trans. Medical Imaging3
2023 The Internet of Sounds: Convergent Trends, Insights, and Future Directions
abstract
Current sound-based practices and systems developed in both academia and industry point to convergent research trends that bring together the field of Sound and Music Computing with that of the Internet of Things. This paper proposes a vision for the emerging field of the Internet of Sounds (IoS), which stems from such disciplines. The IoS relates to the network of Sound Things, i.e., devices capable of sensing, acquiring, processing, actuating, and exchanging data serving the purpose of communicating sound-related information. In the IoS paradigm, which merges under a unique umbrella the emerging fields of the Internet of Musical Things and the Internet of Audio Things, heterogeneous devices dedicated to musical and non-musical tasks can interact and cooperate with one another and with other things connected to the Internet to facilitate sound-based services and applications that are globally available to the users. We survey the state of the art in this space, discuss the technological and non-technological challenges ahead of us and propose a comprehensive research agenda for the field.
Luca Turchet, Mathieu Lagrange, Cristina Rottondi, György Fazekas, Nils Peters, Jan Østergaard, Frederic Font, Tom Bäckström, Carlo Fischione
IEEE Internet Things J.5
2012 Name that room: room identification using acoustic features in a recording
abstract
This paper presents a system for identifying the room in an audio or video recording through the analysis of acoustical properties. The room identification system was tested using a corpus of 13440 reverberant audio samples. With no common content between the training and testing data, an accuracy of 61% for musical signals and 85% for speech signals was achieved. This approach could be applied in a variety of scenarios where knowledge about the acoustical environment is desired, such as location estimation, music recommendation, or emergency response systems.
Nils Peters, Howard Lei, Gerald Friedland
ACM Multimedia1