EDBT 2026 Demo / reviewers in the wild / expert
Karla Schäfer
dblp:356/8208
· DBLP profile ↗
14ranked-venue papers
11as first author
14since 2021 · last 2026
0009-0004-1731-7925ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MOS in Current Voice Conversion Research: Usability and Interpretabilityabstract838 Karla Schäfer, Jeong-Eun Choi |
COMPSAC | 1 |
| 2026 | Beyond Binary: Multi-Class Classification of Audio Deepfakes by TTS and Vocoder Typesabstract2375 Karla Schäfer, Timo Straßburger |
COMPSAC | 1 |
| 2026 | Text Vs. Speech? Detecting Audio Deepfakes on Instagram
Karla Schäfer |
ECIR (2) | 1 |
| 2026 | Too Simple or Too Complex? Using Linguistic Signatures for AI-Generated Text Detection
Karla Schäfer, Mareike Bassenge |
ICPR (8) | 1 |
| 2025 | Disinformation Analysis on Telegram: A Metadata-Centered, Privacy-Aware Datasetabstract2298 Jeong-Eun Choi, Karla Schäfer, York Yannikos, Martin Steinebach |
IEEE Big Data | 2 |
| 2025 | AI-Generated Text Detection Using RoBERTa: A Generalizability and Explainability AnalysisabstractWith the rise of AI-generated text, the need for efficient detectors that perform well on various kinds of text generated by different models and prompts is increasing. We trained and evaluated three detectors: fine-tuned RoBERTa, trained an RoBERTa based adapter and applied adapter fusion on an AI-generated text classification task. The detectors are tested for generalisation on three unseen datasets containing various generation models, generation types, and text styles. All three detectors performed well on the three test sets, outperforming the baselines introduced with the test sets. We found that completely AI-generated text is easier to detect than text that has been manipulated, paraphrased, or rewritten. In contrast, text of less than 50 words was harder to detect than longer text. While texts generated through translating, paraphrasing, and rewriting were better recognized by adapter fusion in most settings; fine-tuned RoBERTa yielded the best overall results. Using transformers-interpret as explainability method and POS-tagging, conjunctions (CC) were identified as characteristics of AI-generated text, whereby personal pronouns (PRP), verbs (present tense; VBP) and modals (MD) were identified as indicators for human-written text, independent of generation models and detectors viewed. Karla Schäfer, Martin Steinebach |
COMPSAC | 1 |
| 2025 | Generalization in Audio Deepfake Detection: Evaluating ASR Encoder-Based Feature ExtractionabstractAudio deepfakes are artificially generated audio recordings that are used to create inauthentic imitations of a person's utterances. In the context of the increasing utilisation of artificial intelligence, the recognition of these recordings is becoming increasingly important. Thereby, the generalization of the detectors is a major challenge in audio deepfake detection (ADD). We analysed six feature representations: MFCC, LFCC, Wav2Vec2.0 and three pre-trained encoders of automatic speech recognition (ASR) systems (Whisper, SpeechT5, Canary) for ADD. For evaluating their generalizability, we trained the models on the ASVspoof 2019 LA train set and tested on the in-the-wild (ITW) test set. Furthermore, we evaluated the training time of the models, analysed the effect of additional (newer) training data and performed a correlation analysis. The ASR encoders were outperformed by using simple MFCC features, achieving an EER of 17.96% on the ITW test set with the lowest training time (7 min. per epoch). The best results were calculated using Wav2Vec 2.0 (+MFCC) with an EER of 17.60%, but using additional training data, and a training time of 135 min. per epoch. Wav2Vec also showed the highest correlations in varying training settings, indicating the best predictability of the model's performance. Karla Schäfer, Martin Steinebach |
DSAA | 1 |
| 2025 | The Sound of Language: A Bilingual Analysis of Voice Conversion and Text-to-Speech SynthesisabstractWith the rise of audio deepfakes, there is an increasing need for comprehensive studies on their generation methods, especially regarding their quality. Areas such as languages beyond English and Chinese, as well as comparisons between voice conversion (VC) and text-to-speech synthesis (TTS), remain underexplored. In our study, we generated samples in English and German using 10 recent VC and TTS methods, including two publicly accessible online tools. We compared these samples using various evaluation methods to gain insights into their quality across different factors. Our analysis indicates that TTS performs slightly better than VC, with minor differences between English and German data. Interestingly, in VC, the gender of the source speaker has minimal influence on the generated samples. Instead, the cross-gender factor appears to affect VC. For both VC and TTS, the target speaker samples used for generation seem to influence the quality of the generated samples. Jeong-Eun Choi, Karla Schäfer, Martin Steinebach |
ICASSP | 2 |
| 2025 | Real-World Audio Deepfake Detection Using SSL-Based Speech Models and Diverse Training DataabstractThe potential for audio deepfakes to be used for malevolent purposes is increasing in line with advances in artificial intelligence and synthesis methods. With this, the need for reliable audio deepfake detectors increased. Most audio deepfake detectors comprise two principal components: a frontend, which is responsible for feature extraction, and a back-end, which performs the classification. Self-supervised learning (SSL) based front-ends are right now the most promising when faced with real-world data. We tested different combinations of six SSL-based front-ends and four back-ends, i.e. classifiers, using nine variously combined training sets, enabling the inclusion of the majority of the currently available training sets for audio deepfake detection. The combination of Wav2Vec2.0 XLS-R (2b) as the front-end and GF as the back-end performed best with an EER of 0.73% on the in-the-wild dataset, outperforming the current SOTA. Furthermore, our findings highlighted the significance of training set constellations and the utilisation of large front-ends. Karla Schäfer, Matthias Neu |
ICTAI | 1 |
| 2025 | AI Got Your Tongue? Analysing the Sounds of Audio Deepfake Generation MethodsabstractIn current research, audio deepfake detectors are trained on finding differences between bona-fide and spoofed samples. A variety of generation methods, mostly distinguished in voice conversion(VC) and text-to-speech synthesis (TTS) exist. We assume that these generation methods lead to specific artefacts in the generated recording. To test this, we created a test set with various spoofs containing the same linguistic content and target speakers as a bona-fide counterpart, using four VC and four TTS models. We applied feature representation methods to compare the differences in 1) bona-fide vs. spoofed samples, 2) samples created using VC vs. TTS and 3) differences in the generation methods used. We found differences between spoofs and bona-fide. Spoofs having overall higher deflections in the waveform and overall smaller values in the spectral evaluation. In the spectral domain, several differences between VC and TTS were detected. XTTS and kNN-VC stood out when viewing the spectral features, e.g. spectral contrast. The samples created using RVC seemed to be the most similar to bona-fide. MFCC and LFCC were the most effective at identifying the differences between bona-fide and spoof audio, making them a suitable choice for detecting audio deepfakes. Karla Schäfer |
ICMR | 1 |
| 2025 | When Voices Deceive: Evaluating and Improving the Robustness of Audio Deepfake Detectors Under Adversarial AttacksabstractLike humans, who are susceptible to illusions, adversarial attacks can deceive audio deepfake detectors into making incorrect judgements. We trained and tested various detection tools against different types of adversarial attack that have been developed in recent years. Previous studies have used the equal error rate to evaluate the effectiveness of a given attack. As it is important to determine whether an attack is effective at tricking the detector into classifying spoofs as bona-fide or vice versa, we used various evaluation metrics to perform a more in-depth analysis. We investigated three adversarial attacks: FGSM, PGDL2 and Malafide. Furthermore, we created an adaptive version of Malafide with filter sizes that change depending on the success of a given attack. The detection models used were LCNN, SpecRNet, RawNet3 and an SSL-based detector consisting of a combination of Wav2Vec2 and AASIST. This selection provides a broad range of detector structures. All detectors were affected by FGSM and PGDL2 in the white box setting; LCNN and SpecRNet the most. Malafide led to a deterioration in white and black box settings, fooling LCNN and SpecRNet into classifying spoofs as bona-fide. Conversely, RawNet3 and Wav2Vec2+AASIST were tricked by Malafide into classifying bona-fide samples as spoofs, resulting in a high number of false alarms while still correctly identifying spoofs. Both models were affected by the filter size; higher-quality degradation in the samples led to worse results, with Wav2Vec2+AASIST being more affected than RawNet3. Through adaptive adversarial training, the detectors partially improved their performance while maintaining their performance on the original data. Karla Schäfer, Leon Ludwig |
TrustCom | 1 |
| 2025 | Machine Learning-Based Detection of AI-Generated Text via Stylistic and Statistical Feature ModelingabstractThrough the advances of large-language models (LLMs) AI- generated text can be created with ease. But, these tools can also pose a threat, e.g. through the creation of disinformation. In this work, we analysed texts generated by three LLMs: GPT-3.5, LLaMA3, and Qwen from the CUDRT dataset. We extracted 220 stylistic and statistical features of human and AI-generated text using the LFTK library. First, we analysed the features using the pearson correlation. Second, we trained five machine learning models and tested the classifiers on detecting completely AI-generated, polished, rewritten texts, and summaries created by AI. We calculated an F1-score of 90%+ for the text generated entirely by AI, depending on the LLM used. We found that AI-generated texts, independent of LLM, can be identified through a high kuperman age, i.e. high word complexity, whereby human-written texts are written with higher lexical variation and richness. We provide an explanation for the classification results and a comparison with RoBERTa (fine-tuned). Karla Schäfer, Martin Steinebach |
TrustCom | 1 |
| 2024 | Comparative Analysis of Voice Conversion in German
Karla Schäfer, Jeong-Eun Choi, Martin Steinebach |
ICPR (22) | 1 |
| 2024 | Scientific Appearance in TelegramabstractThis paper examines the influence of scientific appearance (SA) on post dissemination and analyses a dataset of important actors in Germany, specifically those involved in the dissemination of disinformation on the social media platform Telegram. SA is identified through textual elements such as predefined keywords or digital object identifiers (DOIs). Characteristics and behaviours of actors with and without SA are compared using metadata such as forward counts and original posts. The additional content analysis provides insights into SA's usage and impact. The findings indicate that SA may influence the dissemination of posts and demonstrate how different methods can be applied for studying social media platforms. Jeong-Eun Choi, Karla Schäfer, York Yannikos |
ICWSM | 2 |