EDBT 2026 Demo / reviewers in the wild / expert
Nicolas M. Müller
dblp:248/0679 · also Nicolas Michael Müller
· DBLP profile ↗
15ranked-venue papers
12as first author
13since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 11 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 10 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Easy, Interpretable, Effective: openSMILE for voice deepfake detectionabstractIn this paper, we demonstrate that attacks in the latest ASVspoof5 dataset—a de facto standard in the field of voice authenticity and deepfake detection—can be identified with surprising accuracy using a small subset of very simplistic features. These are derived from the openSMILE library, and are scalar-valued, easy to compute, and human interpretable. For example, attack A10’s unvoiced segments have a mean length of 0.09 ± 0.02, while bona fide instances have a mean length of 0.18 ± 0.07. Using this feature alone, a threshold classifier achieves an Equal Error Rate (EER) of 10.3% for attack A10. Similarly, across all attacks, we achieve up to 0.8% EER, with an overall EER of 15.7 ± 6.0%.We explore the generalization capabilities of these features and find that some of them transfer effectively between attacks, primarily when the attacks originate from similar Text-to-Speech (TTS) architectures. This finding may indicate that voice anti-spoofing is, in part, a problem of identifying and remembering signatures or fingerprints of individual TTS systems. This allows to better understand anti-spoofing models and their challenges in real-world application. Octavian Pascu, Dan Oneata, Horia Cucu, Nicolas M. Müller |
ICASSP | 4 |
| 2025 | Unmasking real-world audio deepfakes: A data-centric approachabstract5343 David Combei, Adriana Cornelia Stan, Dan Oneata, Nicolas M. Müller, Horia Cucu |
INTERSPEECH | 4 |
| 2025 | Replay Attacks Against Audio Deepfake Detectionabstract2245 Nicolas M. Müller, Piotr Kawa, Wei Herng Choong, Adriana Cornelia Stan, Aditya Tirumala Bukkapatnam, Karla Pizzi, Alexander Wagner, Philip Sperl |
INTERSPEECH | 1 |
| 2024 | MLAAD: The Multi-Language Audio Anti-Spoofing DatasetabstractText-to-Speech (TTS) technology brings significant advantages, such as giving a voice to those with speech impairments, but also enables audio deepfakes and spoofs. The former mislead individuals and may propagate misinformation, while the latter undermine voice biometric security systems. AI-based detection can help to address these challenges by automatically differentiating between genuine and fabricated voice recordings. However, these models are only as good as their training data, which currently is severely limited due to an overwhelming concentration on English and Chinese audio in anti-spoofing databases, thus restricting its worldwide effectiveness.In response, this paper presents the Multi-Language Audio Anti-Spoof Dataset (MLAAD), created using 52 TTS models, comprising 22 different architectures, to generate 160.2 hours of synthetic voice in 23 different languages. We train and evaluate three state-of-the-art deepfake detection models with MLAAD, and observe that MLAAD demonstrates superior performance over comparable datasets like InTheWild or FakeOrReal when used as a training resource. Furthermore, in comparison with the renowned ASVspoof 2019 dataset, MLAAD proves to be a complementary resource. In tests across eight datasets, MLAAD and ASVspoof 2019 alternately outperformed each other, both excelling on four datasets.By publishing1MLAAD and making trained models accessible via an interactive webserver2, we aim to democratize antispoofing technology, making it accessible beyond the realm of specialists, thus contributing to global efforts against audio spoofing and deepfakes. Nicolas M. Müller, Piotr Kawa, Wei Herng Choong, Edresson Casanova, Eren Gölge, Piotr Syga, Philip Sperl, Konstantin Böttinger |
IJCNN | 1 |
| 2024 | Shortcut Detection With Variational AutoencodersabstractIn practical machine learning (ML) applications, it is vital for models to make predictions based on robust, generalizable features rather than unreliable data patterns. For example, in supervised classification tasks, a model may mistakenly label an image as ‘horse’ not due to identifying the animal’s traits, but because of a recurring watermark in ‘horse’ images — a learning shortcut. These deceptive shortcuts lead to artificially high performance in training and testing, creating a misleading impression of the model’s actual effectiveness. Such models often fail in real-world scenarios when these accidental correlations are absent. Thus, identifying and addressing these spurious correlations is a critical yet underexplored challenge.In our study, we introduce a new method for detecting such shortcuts in image and audio datasets. We use variational autoencoders (VAE) to separate features in the latent space of the VAE. This enables clear and semi-automatic identification of feature-target correlations in datasets. Our approach’s effectiveness is demonstrated on various real-world datasets, uncovering previously undetected shortcuts. For instance, we find that in fruit classification, the class prediction is influenced chiefly by the camera’s distance from the fruit.Our approach not only sheds light on the intricacies of what machine learning models learn but also aids in circumventing unwanted correlations that might limit their practical effectiveness. The tool is open-source and can be accessed at ANONYMOUS_URL. Nicolas M. Müller, Simon Roschmann, Shahbaz Farooque Khan, Philip Sperl, Konstantin Böttinger |
IJCNN | 1 |
| 2024 | Harder or Different? Understanding Generalization of Audio Deepfake Detectionabstract2705 Nicolas M. Müller, Nicholas W. D. Evans, Hemlata Tak, Philip Sperl, Konstantin Böttinger |
INTERSPEECH | 1 |
| 2024 | A New Approach to Voice Authenticityabstract2245 Nicolas M. Müller, Piotr Kawa, Shen Hu, Matthias Neu, Jennifer Williams 0001, Philip Sperl, Konstantin Böttinger |
INTERSPEECH | 1 |
| 2023 | Protecting Publicly Available Data With Machine Learning Shortcuts
Nicolas M. Müller, Maximilian Burgert, Pascal Debus, Jennifer Williams 0001, Philip Sperl, Konstantin Böttinger |
BMVC | 1 |
| 2023 | Complex-valued neural networks for voice anti-spoofingabstract3814 Nicolas M. Müller, Philip Sperl, Konstantin Böttinger |
INTERSPEECH | 1 |
| 2022 | Does Audio Deepfake Detection Generalize?abstract2783 Nicolas M. Müller, Pavel Czempin, Franziska Dieckmann, Adam Froghyar, Konstantin Böttinger |
INTERSPEECH | 1 |
| 2022 | Attacker Attribution of Audio Deepfakesabstract2788 Nicolas M. Müller, Franziska Dieckmann, Jennifer Williams 0001 |
INTERSPEECH | 1 |
| 2021 | Adversarial Vulnerability of Active Transfer Learning
Nicolas M. Müller, Konstantin Böttinger |
IDA | 1 |
| 2021 | SC-GlowTTS: An Efficient Zero-Shot Multi-Speaker Text-To-Speech ModelabstractIn this paper, we propose SC-GlowTTS: an efficient zero-shot multi-speaker text-to-speech model that improves similarity for speakers unseen during training. We propose a speaker-conditional architecture that explores a flow-based decoder that works in a zero-shot scenario. As text encoders, we explore a dilated residual convolutional-based encoder, gated convolutional-based encoder, and transformer-based encoder. Additionally, we have shown that adjusting a GAN-based vocoder for the spectrograms predicted by the TTS model on the training dataset can significantly improve the similarity and speech quality for new speakers. Our model converges using only 11 speakers, reaching state-of-the-art results for similarity with new speakers, as well as high speech quality. Edresson Casanova, Christopher Shulby, Eren Gölge, Nicolas M. Müller, Frederico Santos de Oliveira, Arnaldo Cândido Jr., Anderson da Silva Soares, Sandra M. Aluísio, Moacir Ponti |
Interspeech | 4 |
| 2020 | Data Poisoning Attacks on Regression Learning and Corresponding DefensesabstractAdversarial data poisoning is an effective attack against machine learning and threatens model integrity by introducing poisoned data into the training dataset. So far, it has been studied mostly for classification, even though regression learning is used in many mission critical systems (such as dosage of medication, control of cyber-physical systems and managing power supply). Therefore, in the present research, we aim to evaluate all aspects of data poisoning attacks on regression learning, exceeding previous work both in terms of breadth and depth. We present realistic scenarios in which data poisoning attacks threaten production systems and introduce a novel black-box attack, which is then applied to a real-word medical use-case. As a result, we observe that the mean squared error (MSE) of the regressor increases to 150 percent due to inserting only two percent of poison samples. Finally, we present a new defense strategy against the novel and previous attacks and evaluate it thoroughly on 26 datasets. As a result of the conducted experiments, we conclude that the proposed defence strategy effectively mitigates the considered attacks. Nicolas M. Müller, Daniel Kowatsch, Konstantin Böttinger |
PRDC | 1 |
| 2019 | Identifying Mislabeled Instances in Classification DatasetsabstractA key requirement for supervised machine learning is labeled training data, which is created by annotating unlabeled data with the appropriate class. Because this process can in many cases not be done by machines, labeling needs to be performed by human domain experts. This process tends to be expensive both in time and money, and is prone to errors. Additionally, reviewing an entire labeled dataset manually is often prohibitively costly, so many real world datasets contain mislabeled instances.To address this issue, we present in this paper a non-parametric end-to-end pipeline to find mislabeled instances in numerical, image and natural language datasets. We evaluate our system quantitatively by adding a small number of label noise to 29 datasets, and show that we find mislabeled instances with an average precision of more than 0.84 when reviewing our system's top 1% recommendation. We then apply our system to publicly available datasets and find mislabeled instances in CIFAR-100, Fashion-MNIST, and others. Finally, we publish the code and an applicable implementation of our approach. Nicolas M. Müller, Karla Markert |
IJCNN | 1 |