VLDB 2026 Research / reviewers in the wild / expert
Manuel Milling
dblp:255/0337
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-8842-2958ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Discourse Features Enhance Detection of Document-Level Machine-Generated ContentabstractThe availability of high-quality APIs for Large Language Models (LLMs) has facilitated the widespread creation of Machine-Generated Content (MGC), posing challenges such as academic plagiarism and the spread of misinformation. Existing MGC detectors often focus solely on surface-level information, overlooking implicit and structural features. This makes them susceptible to deception by surface-level sentence patterns, particularly for longer texts and in texts that have been subsequently paraphrased. To overcome these challenges, we introduce novel methodologies and datasets. Besides the publicly available dataset Plagbench, we developed the paraphrased Long-Form Question and Answer (paraLFQA) and paraphrased Writing Prompts (paraWP) datasets using GPT and DIPPER, a discourse paraphrasing tool, by extending artifacts from their original versions. To better capture the structure of longer texts at document level, we propose DTransformer, a model that integrates discourse analysis through PDTB preprocessing to encode structural features. It results in substantial performance gains across both datasets – 15.5% absolute improvement on paraLFQA, 4% absolute improvement on paraWP, and 1.5% absolute improvemene on M4 compared to SOTA approaches. The data and code are available at this link1. Yupei Li, Manuel Milling, Lucia Specia, Björn W. Schuller |
IJCNN | 2 |
| 2025 | Audio-Based Kinship Verification Using Age Domain ConversionabstractAudio-based kinship verification (AKV) is important in many domains, such as home security monitoring, forensic identification, and social network analysis. A key challenge in the task arises from differences in age across samples from different individuals, which can be interpreted as a domain bias in a cross-domain verification task. To address this issue, we design the notion of an “age-standardised domain” wherein we utilise the optimised CycleGAN-VC3 network to perform age-audio conversion to generate the in-domain audio. The generated audio dataset is employed to extract a range of features, which are then fed into a metric learning architecture to verify kinship. Experiments are conducted on the KAN_AV audio dataset.The results demonstrate that the method markedly enhances the accuracy of kinship verification, while also offering novel insights for future kinship verification research. Alican Akman, Xin Jing 0001, Manuel Milling, Björn W. Schuller |
IEEE Signal Process. Lett. | 4 |
| 2024 | Bringing the Discussion of Minima Sharpness to the Audio Domain: A Filter-Normalised Evaluation for Acoustic Scene ClassificationabstractThe correlation between the sharpness of loss minima and generalisation in the context of deep neural networks has been subject to discussion for a long time. Whilst mostly investigated in the context of selected benchmark data sets in the area of computer vision, we explore this aspect for the acoustic scene classification task of the DCASE2020 challenge data. Our analysis is based on two-dimensional filter-normalised visualisations and a derived sharpness measure. Our exploratory analysis shows that sharper minima tend to show better generalisation than flat minima –even more so for out-of-domain data, recorded from previously unseen devices–, thus adding to the dispute about better generalisation capabilities of flat minima. We further find that, in particular, the choice of optimisers is a main driver of the sharpness of minima and we discuss resulting limitations with respect to comparability. Our code, trained model states and loss landscape visualisations are publicly available. Manuel Milling, Andreas Triantafyllopoulos, Iosif Tsangko, Simon David Noel Rampp, Björn W. Schuller |
ICASSP | 1 |
| 2024 | Are you sure? Analysing Uncertainty Quantification Approaches for Real-world Speech Emotion Recognition
Oliver Schrüfer, Manuel Milling, Felix Burkhardt, Florian Eyben, Björn W. Schuller |
INTERSPEECH | 2 |
| 2024 | INTERSPEECH 2009 Emotion Challenge Revisited: Benchmarking 15 Years of Progress in Speech Emotion Recognition
Andreas Triantafyllopoulos, Anton Batliner, Simon David Noel Rampp, Manuel Milling, Björn W. Schuller |
INTERSPEECH | 4 |
| 2024 | Automatic Bird Sound Source Separation Based on Passive Acoustic Devices in Wild EnvironmentabstractThe Internet of Things (IoT)-based passive acoustic monitoring (PAM) has shown great potential in large-scale remote bird monitoring. However, field recordings often contain overlapping signals, making precise bird information extraction challenging. To solve this challenge, first, the inter-channel spatial feature is chosen as complementary information to the spectral feature to obtain additional spatial correlations between the sources. Then, an end-to-end model named BACPPNet is built based on Deeplabv3plus and enhanced with the polarized self-attention mechanism to estimate the spectral amplitude mask (SMM) for separating bird vocalizations. Finally, the separated bird vocalizations are recovered from SMMs and the spectrogram of mixed audio using the inverse short Fourier transform (ISTFT). We evaluate our proposed method utilizing the generated mixed dataset. Experiments have shown that our method can separate bird vocalizations from mixed audio with RMSE, SDR, SIR, SAR, and STOI values of 2.82, 10.00dB, 29.90 dB, 11.08 dB, and 0.66, respectively, which are better than existing methods. Furthermore, the average classification accuracy of the separated bird vocalizations drops the least. This indicates that our method outperforms other compared separation methods in bird sound separation and preserves the fidelity of the separated sound sources, which might help us better understand wild bird sound recordings. Jiangjian Xie, Yuwei Shi, Dongming Ni, Manuel Milling, Shuo Liu 0012, Junguo Zhang, Kun Qian 0003, Björn W. Schuller |
IEEE Internet Things J. | 4 |
| 2024 | Audio Enhancement for Computer Audition - An Iterative Training Paradigm Using Sample Importance
Manuel Milling, Shuo Liu 0012, Andreas Triantafyllopoulos, Ilhan Aslan, Björn W. Schuller |
J. Comput. Sci. Technol. | 1 |
| 2021 | A Prototypical Network Approach for Evaluating Generated Emotional SpeechabstractThe collection of emotional speech data is a time-consuming and costly endeavour.Generative networks can be applied to augment the limited audio data artificially.However, it is challenging to evaluate generated audio for its similarity to source data, as current quantitative metrics are not necessarily suited to the audio domain.We explore the use of a prototypical network to evaluate four classes of generated emotional audio with this in mind.We first extract spectrogram images from WAVEGAN generated audio and other audio augmentation approaches, comparing similarity to the class prototype and diversity within the embedding space.Furthermore, we augment the source training set with each augmentation type and perform a classification to explore the generated audio plausibility.Results suggest that quality and diversity can be quantitatively observed with this approach.In the chosen context, we see that WAVEGAN generated data is recognisable as a source data class (F1-score 43.6 %), and the samples add similar diversity as unseen source data.This result leads to more plausible data for augmentation of the source training set -achieving up to 63.9 % F1 which is a 3.5 % improvement over the source data baseline. Alice Baird, Silvan Mertes, Manuel Milling, Lukas Stappen, Thomas Wiest, Elisabeth André, Björn W. Schuller |
Interspeech | 3 |
| 2021 | Emotion Recognition in Public Speaking Scenarios Utilising An LSTM-RNN Approach with AttentionabstractSpeaking in public can be a cause of fear for many people. Research suggests that there are physical markers such as an increased heart rate and vocal tremolo that indicate an individual's state of wellbeing during a public speech. In this study, we explore the advantages of speech-based features for continuous recognition of the emotional dimensions of arousal and valence during a public speaking scenario. Furthermore, we explore biological signal fusion, and perform cross-language (German and English) analysis by training language-independent models and testing them on speech from various native and non-native speaker groupings. For the emotion recognition task itself, we utilise a Long Short-Term Memory - Recurrent Neural Network (LSTM-RNN) architecture with a self-attention layer. When utilising audio-only features and testing with non-native German's speaking German we achieve at best a concordance correlation coefficient (CCC) of 0.640 and 0.491 for arousal and valence, respectively - demonstrating a strong effect for this task from non-native speakers, as well as promise for the suitability of deep learning for continuous emotion recognition in the context of public speaking. Alice Baird, Shahin Amiriparian, Manuel Milling, Björn W. Schuller |
SLT | 3 |