Tim Polzehl

dblp:05/7474 · DBLP profile ↗
← Back
33ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0001-9592-0296ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 22 · 6 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 4Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Avatar Motion Signatures: Evaluating Linkability of Expressive De-Identification
Fenja Schulz, Jan Marquenie, Carlos Franzreb, Tim Polzehl, Ingo Siegert, Sebastian Möller 0001
ICISSP (2)4
2025 On the Effect of Dataset Size and Composition for Privacy Evaluation
Danai Georgiou, Carlos Franzreb, Tim Polzehl
ICISSP (2)3
2025 Generalizable Audio Spoofing Detection using Non-Semantic Representations
abstract
Rapid advancements in generative modeling have made synthetic audio generation easy, making speech-based services vulnerable to spoofing attacks. Consequently, there is a dire need for robust countermeasures more than ever. Existing solutions for deepfake detection are often criticized for lacking generalizability and fail drastically when applied to real-world data. This study proposes a novel method for generalizable spoofing detection leveraging non-semantic universal audio representations. Extensive experiments have been performed to find suitable non-semantic features using TRILL and TRILLsson models. The results indicate that the proposed method achieves comparable performance on the in-domain test set while significantly outperforming state-of-the-art approaches on out-of-domain test sets. Notably, it demonstrates superior generalization on public-domain data, surpassing methods based on hand-crafted features, semantic embeddings, and end-to-end architectures.
Yassine El Kheir, Carlos Franzreb, Tim Herzig, Tim Polzehl, Sebastian Möller 0001
INTERSPEECH5
2025 Private kNN-VC: Interpretable Anonymization of Converted Speech
Carlos Franzreb, Tim Polzehl, Sebastian Möller 0001
INTERSPEECH3
2025 BiCrossMamba-ST: Speech Deepfake Detection with Bidirectional Mamba Spectro-Temporal Cross-Attention
Yassine El Kheir, Tim Polzehl, Sebastian Möller 0001
INTERSPEECH2
2024 Anonymising Elderly and Pathological Speech: Voice Conversion Using DDSP and Query-by-Example
abstract
Speech anonymisation aims to protect speaker identity by changing personal identifiers in speech while retaining linguistic content. Current methods fail to retain prosody and unique speech patterns found in elderly and pathological speech domains, which is essential for remote health monitoring. To address this gap, we propose a voice conversion-based method (DDSP-QbE) using differentiable digital signal processing and query-by-example. The proposed method, trained with novel losses, aids in disentangling linguistic, prosodic, and domain representations, enabling the model to adapt to uncommon speech patterns. Objective and subjective evaluations show that DDSP-QbE significantly outperforms the voice conversion state-of-the-art concerning intelligibility, prosody, and domain preservation across diverse datasets, pathologies, and speakers while maintaining quality and speaker anonymity. Experts validate domain preservation by analysing twelve clinically pertinent domain attributes.
Suhita Ghosh, Mélanie Jouaiti, Yamini Sinha, Tim Polzehl, Ingo Siegert, Sebastian Stober
INTERSPEECH5
2024 Towards Classifying Mother Tongue from Infant Cries - Findings Substantiating Prenatal Learning Theory
Tim Polzehl, Tim Herzig, Friedrich Wicke, Kathleen Wermke, Razieh Khamsehashari, Michiko Dahlem, Sebastian Möller 0001
INTERSPEECH1
2023 Emo-StarGAN: A Semi-Supervised Any-to-Many Non-Parallel Emotion-Preserving Voice Conversion
abstract
Speech anonymisation prevents misuse of spoken data by removing any personal identifier while preserving at least linguistic content. However, emotion preservation is crucial for natural human-computer interaction. The well-known voice conversion technique StarGANv2-VC achieves anonymisation but fails to preserve emotion. This work presents an any-to-many semi-supervised StarGANv2-VC variant trained on partially emotion-labelled non-parallel data. We propose emotion-aware losses computed on the emotion embeddings and acoustic features correlated to emotion. Additionally, we use an emotion classifier to provide direct emotion supervision. Objective and subjective evaluations show that the proposed approach significantly improves emotion preservation over the vanilla StarGANv2-VC. This considerable improvement is seen over diverse datasets, emotions, target speakers, and inter-group conversions without compromising intelligibility and anonymisation.
Suhita Ghosh, Yamini Sinha, Ingo Siegert, Tim Polzehl, Sebastian Stober
INTERSPEECH5
2022 Speaker adaptation for Wav2vec2 based dysarthric ASR
Murali Karthick Baskar, Tim Herzig, Diana Nguyen, Mireia Díez, Tim Polzehl, Lukás Burget, Jan Cernocký
INTERSPEECH5
2022 Towards Automated Dialog Personalization using MBTI Personality Indicators
Daniel Fernau, Stefan Hillmann, Nils Feldhus, Tim Polzehl
INTERSPEECH4
2022 Towards Personality-Aware Chatbots
abstract
Chatbots are increasingly used to automate operational processes in customer service.However, most chatbots lack adaptation towards their users which may results in an unsatisfactory experience.Since knowing and meeting personal preferences is a key factor for enhancing usability in conversational agents, in this study we analyze an adaptive conversational agent that can automatically adjust according to a user's personality type carefully excerpted from the Myers-Briggs type indicators.An experiment including 300 crowd workers examined how typifications like extroversion/introversion and thinking/feeling can be assessed and designed for a conversational agent in a job recommender domain.Our results validate the proposed design choices, and experiments on a user-matched personality typification, following the so-called law of attraction rule, show a significant positive influence on a range of selected usability criteria such as overall satisfaction, naturalness, promoter score, trust and appropriateness of the conversation.
Daniel Fernau, Stefan Hillmann, Nils Feldhus, Tim Polzehl, Sebastian Möller 0001
SIGDIAL4
2021 Argument Mining in Tweets: Comparing Crowd and Expert Annotations for Automated Claim and Evidence Detection
Neslihan Iskender, Robin Schaefer, Tim Polzehl, Sebastian Möller 0001
NLDB3
2020 Towards a Reliable and Robust Methodology for Crowd-Based Subjective Quality Assessment of Query-Based Extractive Text Summarization
abstract
The intrinsic and extrinsic quality evaluation is an essential part of the summary evaluation methodology usually conducted in a traditional controlled laboratory environment. However, processing large text corpora using these methods reveals expensive from both the organizational and the financial perspective. For the first time, and as a fast, scalable, and cost-effective alternative, we propose micro-task crowdsourcing to evaluate both the intrinsic and extrinsic quality of query-based extractive text summaries. To investigate the appropriateness of crowdsourcing for this task, we conduct intensive comparative crowdsourcing and laboratory experiments, evaluating nine extrinsic and intrinsic quality measures on 5-point MOS scales. Correlating results of crowd and laboratory ratings reveals high applicability of crowdsourcing for the factors overall quality, grammaticality, non-redundancy, referential clarity, focus, structure & coherence, summary usefulness, and summary informativeness. Further, we investigate the effect of the number of repetitions of assessments on the robustness of mean opinion score of crowd ratings, measured against the increase of correlation coefficients between crowd and laboratory. Our results suggest that the optimal number of repetitions in crowdsourcing setups, in which any additional repetitions do no longer cause an adequate increase of overall correlation coefficients, lies between seven and nine for intrinsic and extrinsic quality factors.
Neslihan Iskender, Tim Polzehl, Sebastian Möller 0001
LREC2
2020 Are You Still Watching? Streaming Video Quality and Engagement Assessment in the Crowd
abstract
As video streaming accounts for the majority of Internet traffic, monitoring its quality is of importance to both Over the Top (OTT) providers as well as Internet Service Providers (ISPs). While OTTs have access to their own analytics data with detailed information, ISPs often have to rely on automated network probes for estimating streaming quality, and likewise, academic researchers have no information on actual customer behavior. In this paper, we present first results from a large-scale crowdsourcing study in which three major video streaming OTTs were compared across five major national ISPs in Germany. We not only look at streaming performance in terms of loading times and stalling, but also customer behavior (e.g., user engagement) and Quality of Experience based on the ITU-T P.1203 QoE model. We used a browser extension to evaluate the streaming quality and to passively collect anonymous OTT usage information based on explicit user consent. Our data comprises over 400,000 video playbacks from more than 2,000 users, collected throughout the entire year of 2019. The results show differences in how customers use the video services, how the content is watched, how the network influences video streaming QoE, and how user engagement varies by service. Hence, the crowdsourcing paradigm is a viable approach for third parties to obtain streaming QoE insights from OTTs.
Werner Robitza, Alexander M. Dethof, Steve Goering, Alexander Raake, André Beyer, Tim Polzehl
QoMEX6
2019 A Crowdsourcing Approach to Evaluate the Quality of Query-based Extractive Text Summaries
abstract
High cost and time consumption are concurrent barriers for research and application of automated summarization. In order to explore options to overcome this barrier, we analyze the feasibility and appropriateness of micro-task crowdsourcing for evaluation of different summary quality characteristics and report an ongoing work on the crowdsourced evaluation of query-based extractive text summaries. To do so, we assess and evaluate a number of linguistic quality factors such as grammaticality, non-redundancy, referential clarity, focus and structure & coherence. Our first results imply that referential clarity, focus and structure & coherence are the main factors effecting the perceived summary quality by crowdworkers. Further, we compare these results using an initial set of expert annotations that is currently being collected, as well as an initial set of automatic quality score ROUGE for summary evaluation. Preliminary results show that ROUGE does not correlate with linguistic quality factors, regardless if assessed by crowd or experts. Further, crowd and expert ratings show highest degree of correlation when assessing low quality summaries. Assessments increasingly divert when attributing high quality judgments.
Neslihan Iskender, Aleksandra Gabryszak, Tim Polzehl, Leonhard Hennig, Sebastian Möller 0001
QoMEX3
2016 Crowdsourcing a Multi-lingual Speech Corpus: Recording, Transcription and Annotation of the CrowdIS Corpora
Andrew Caines, Christian Bentz, Calbert Graham, Tim Polzehl, Paula Buttery
LREC4
2015 Effect of trapping questions on the reliability of speech quality judgments in a crowdsourcing paradigm
Babak Naderi, Tim Polzehl, Ina Wechsung, Friedemann Köster, Sebastian Möller 0001
INTERSPEECH2
2015 Advanced crowdsourcing for speech and beyond: introduction by the organizers
Tim Polzehl, Gina-Anne Levow
INTERSPEECH1
2015 Robustness in speech quality assessment and temporal training expiry in mobile crowdsourcing environments
Tim Polzehl, Babak Naderi, Friedemann Köster, Sebastian Möller 0001
INTERSPEECH1
2014 Crowdee: mobile crowdsourcing micro-task platform for celebrating the diversity of languages
Babak Naderi, Tim Polzehl, André Beyer, Tibor Pilz, Sebastian Möller 0001
INTERSPEECH2
2012 Articulatory features for expressive speech synthesis
abstract
This paper describes some of the results from the project entitled “New Parameterization for Emotional Speech Synthesis” held at the Summer 2011 JHU CLSP workshop. We describe experiments on how to use articulatory features as a meaningful intermediate representation for speech synthesis. This parameterization not only allows us to reproduce natural sounding speech but also allows us to generate stylistically varying speech.
Alan W. Black, H. Timothy Bunnell, Ying Dou, Prasanna Kumar Muthukumar, Florian Metze, Daniel Perry 0002, Tim Polzehl, Kishore Prahallad, Stefan Steidl, Callie Vaughn
ICASSP7
2012 On Speaker-Independent Personality Perception and Prediction from Speech
abstract
In this paper, we present ongoing experiments and insights regarding automatic assessment of perceived personality. While within the INTERSPEECH Speaker Trait Challenge participants will train systems in order to recognize binary targets along the Big 5 personality trait, we will analyze and discuss properties of the data, the labeling scheme and the predictive quality. Conducting factor analyses, estimating reliability, and building regression models capturing dimensions of personality we compare all results to our former and current work and introduce a new extension of our personality database. Eventually, this paper contributes in methodology and understanding on how to asses the perceived personality from an unknown speaker by humans and machines.
Tim Polzehl, Katrin Schoenenberg, Sebastian Möller 0001, Florian Metze, Gelareh Mohammadi, Alessandro Vinciarelli
INTERSPEECH1
2011 Modeling Speaker Personality Using Voice
abstract
In this paper, we validate the application of an established personality assessment and modeling paradigm to speech input, and extend earlier work towards text independent speech input. We show that human labelers can consistently label acted speech data generated across multiple recording sessions, and investigate further which of the 5 scales in the NEO-FFI scheme can be assessed from speech, and how a manipulation of one scale influences the perception of another. Finally, we present a clustering of human labels of perceived personality traits, which will be useful in future experiments on automatic classification and generation of personality traits from speech.
Tim Polzehl, Sebastian Möller 0001, Florian Metze
INTERSPEECH1
2011 Anger recognition in speech using acoustic and linguistic cues
Tim Polzehl, Alexander Schmitt, Florian Metze, Michael Wagner 0004
Speech Commun.1
2010 Late fusion of individual engines for improved recognition of negative emotion in speech - learning vs. democratic vote
abstract
The fusion of multiple recognition engines is known to be able to outperform individual ones, given sufficient independence of methods, models, and knowledge sources. We therefore investigate late fusion of different speech-based recognizers of emotion. Two generally different streams of information are considered: acoustics and linguistics fed by state-of-the-art automatic speech recognition. A total of five emotion recognition engines from different sites that provide heterogeneous output information are integrated by either simple democratic vote or learning `which predictor to trust when'. We are able to significantly outperform the best individual engine by fusion, and the so far best reported result on the recently introduced Emotion Challenge task.
Björn W. Schuller, Florian Metze, Stefan Steidl, Anton Batliner, Florian Eyben, Tim Polzehl
ICASSP6
2010 Emotion recognition using imperfect speech recognition
abstract
This paper investigates the use of speech-to-text methods for assigning an emotion class to a given speech utterance. Previous work shows that an emotion extracted from text can convey complementary evidence to the information extracted by classifiers based on spectral, or other non-linguistic features. As speech-to-text usually presents significantly more computational effort, in this study we investigate the degree of speech-to-text accuracy needed for reliable detection of emotions from an automatically generated transcription of an utterance. We evaluate the use of hypotheses in both training and testing, and compare several classification approaches on the same task. Our results show that emotion recognition performance stays roughly constant as long as word accuracy doesn't fall below a reasonable value, making the use of speech-to-text viable for training of emotion classifiers based on linguistics.
Florian Metze, Anton Batliner, Florian Eyben, Tim Polzehl, Björn W. Schuller, Stefan Steidl
INTERSPEECH4
2010 Comparison of approaches for instrumentally predicting the quality of text-to-speech systems
abstract
In this paper, we compare and combine different approaches for instrumentally predicting the perceived quality of Text-to-Speech systems. First, a log-likelihood is determined by comparing features extracted from the synthesized speech signal with features trained on natural speech. Second, parameters are extracted which capture quality-relevant degradations of the synthesized speech signal. Both approaches are combined and evaluated on three auditory test databases. The results show that auditory quality judgments can in many cases be predicted with a sufficiently high accuracy and reliability, but that there are considerable differences, mainly between male and female speech samples.
Sebastian Möller 0001, Florian Hinterleitner, Tiago H. Falk, Tim Polzehl
INTERSPEECH4
2010 The Influence of the Utterance Length on the Recognition of Aged Voices
Alexander Schmitt, Tim Polzehl, Wolfgang Minker, Jackson Liscombe
LREC2
2010 Automatically assessing acoustic manifestations of personality in speech
abstract
In this paper, we present first results on applying a personality assessment paradigm to speech input, and comparing human and automatic performance on this task. We cue a professional speaker to produce speech using different personality profiles and encode the resulting vocal personality impressions in terms of the Big Five NEO-FFI personality traits. We then have human raters, who do not know the speaker, estimate the five factors. We analyze the recordings using signal-based acoustic and prosodic methods and observe high consistency between the acted personalities, the raters' assessments, and initial automatic classification results. This presents a first step towards being able to handle personality traits in speech, which we envision will be used in future voice-based communication between humans and machines.
Tim Polzehl, Sebastian Möller 0001, Florian Metze
SLT1
2009 Fall and emergency detection with mobile phones
abstract
In this demo, we present an application for mobile phones which can monitor physical activities of users and detect unexpected emergency situations such as a sudden fall or accident. Upon detection of such an event, the mobile phone can inform a designated center (by automatically calling or sending message) about the incident and its location. This can facilitate and speed up recovery and help process especially if the user is alone or the accident has happened in a deserted place. Such an application can be particularly useful for elderly people or people with physical and movement disabilities. The application operates based on analysis of user movements using data provided by accelerometers integrated in mobile phones.
Hamed Ketabdar, Tim Polzehl
ASSETS2
2009 Tactile and visual alerts for deaf people by mobile phones
abstract
In this demo, we present an application for mobile phones which can analyse audio context and issue tactile or visual alerts if an audio event happens. This application can be useful especially for deaf or hard of hearing people to be alerted of audio events happening around them. The audio context analysis algorithm captures data using mobile phone's microphone and checks for changes in audio activities around the user. If such a change happens and also some other circumstances are met, the application issues visual or vibro-tactile alerts (by vibrating mobile phone) proportional to the change in audio context. This informs the user about an event. The functionality of this algorithm can be further enhanced by analysis of user movements.
Hamed Ketabdar, Tim Polzehl
ASSETS2
2009 Detecting real life anger
abstract
Acoustic anger detection in voice portals can help to enhance human computer interaction. A comprehensive voice portal data collection has been carried out and gives new insight on the nature of real life data. Manual labeling revealed a high percentage of non-classifiable data. Experiments with a statistical classifier indicate that, in contrast to pitch and energy related features, duration measures do not play an important role for this data while cepstral information does. Also in a direct comparison between Gaussian Mixture Models and Support Vector Machines the latter gave better results.
Felix Burkhardt, Tim Polzehl, Joachim Stegmann, Florian Metze, Richard Huber
ICASSP2
2009 Emotion classification in children's speech using fusion of acoustic and linguistic features
abstract
This paper describes a system to detect angry vs. non-angry utterances of children who are engaged in dialog with an Aibo robot dog. The system was submitted to the Interspeech2009 Emotion Challenge evaluation. The speech data consist of short utterances of the children’s speech, and the proposed system is designed to detect anger in each given chunk. Frame-based cepstral features, prosodic and acoustic features as well as glottal excitation features are extracted automatically, reduced in dimensionality and classified by means of an artificial neural network and a support vector machine. An automatic speech recognizer transcribes the words in an utterance and yields a separate classification based on the degree of emotional salience of the words. Late fusion is applied to make a final decision on anger vs. non-anger of the utterance. Preliminary results show 75.9% unweighted average recall on the training data and 67.6 % on the test set. Index Terms: speech processing, meta-data extraction, emotion recognition, evaluation
Tim Polzehl, Shiva Sundaram, Hamed Ketabdar, Michael Wagner 0004, Florian Metze
INTERSPEECH1