Petra Wagner

dblp:42/9238 · DBLP profile ↗
← Back
56ranked-venue papers
8as first author
15since 2021 · last 2025
0000-0001-6662-3612ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 47 · 8 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Speech Synthesis along Perceptual Voice Quality Dimensions
abstract
While expressive speech synthesis or voice conversion systems mainly focus on controlling or manipulating abstract prosodic characteristics of speech, such as emotion or accent, we here address the control of perceptual voice qualities (PVQs) recognized by phonetic experts, which are speech properties at a lower level of abstraction. The ability to manipulate PVQs can be a valuable tool for teaching speech pathologists in training or voice actors. In this paper, we integrate a Conditional Continuous-Normalizing-Flow-based method into a Text-to-Speech system to modify perceptual voice attributes on a continuous scale. Unlike previous approaches, our system avoids direct manipulation of acoustic correlates and instead learns from examples. We demonstrate the system's capability by manipulating four voice qualities: Roughness, breathiness, resonance and weight. Phonetic experts evaluated these modifications, both for seen and unseen speaker conditions. The results highlight both the system's strengths and areas for improvement.
Frederik Rautenberg, Michael Kuhlmann, Fritz Seebauer, Jana Wiechmann, Petra Wagner, Reinhold Häb-Umbach
ICASSP5
2025 Towards Frame-level Quality Predictions of Synthetic Speech
abstract
Kuhlmann M, Seebauer FM, Wagner P, Haeb-Umbach R. Towards Frame-level Quality Predictions of Synthetic Speech. In: Interspeech 2025. Interspeech. Baixas: International Speech Communication Association; 2025: 2300-2304.
Michael Kuhlmann, Fritz Seebauer, Petra Wagner, Reinhold Häb-Umbach
INTERSPEECH3
2025 Synthesizing Speech with Selected Perceptual Voice Qualities - A Case Study with Creaky Voice
abstract
Rautenberg F, Seebauer FM, Wiechmann J, Kuhlmann M, Wagner P, Haeb-Umbach R. Synthesizing Speech with Selected Perceptual Voice Qualities – A Case Study with Creaky Voice. In: Interspeech 2025. ISCA: ISCA; 2025: 1633-1637.
Frederik Rautenberg, Fritz Seebauer, Jana Wiechmann, Michael Kuhlmann, Petra Wagner, Reinhold Häb-Umbach
INTERSPEECH5
2025 A comparison of discrete and continuous prominence perception methods in German
abstract
In this paper we report on three methods to investigate syllable-based prominence identification on a set of German read sentences: prosodic expert annotations of pitch accentuation, as well as a Rapid Prosody Transcription (RPT) style task and a finger-tapping task performed by naive listeners. In the present study, audio recordings of the speech materials used to elicit prominence judgments were supplemented by signals collected with miniature accelerometers placed on the throat skin below the glottis , allowing for a more reliable investigation of the contribution of voice quality. Various other signal-based parameters are correlated with prominence judgments to confirm if findings from previous work on word-based prominence judgments also hold for judgments at the syllable level. Results replicate past findings for German and other languages: the presence and type of pitch accentuation are reliable predictors of prominence judgments by naive listeners, on both the RPT and tapping task. Various individual acoustic parameters such as f0, duration and intensity were once again found to systematically covary with greater perceived prominence. A direct comparison of tapping and RPT results moreover indicated that beyond pitch accent related factors, listeners may employ different strategies to judge prominence as a function of task. In the RPT, they rely more on their knowledge of whether a given syllable carries lexical stress, whereas in the tapping task, they attend relatively strongly to acoustic duration. It is also shown that voice quality varies along with prominence ratings, but less strongly than other features such as duration.
Anna Bruggeman, Marcin Wlodarczak, Petra Wagner
Speech Commun.3
2024 Predictability of Understanding in Explanatory Interactions Based on Multimodal Cues
abstract
In explanatory interactions, explainees are expected to continuously provide feedback to explainers by signaling whether they understand an ongoing explanation. The study presented in this paper is based on the hypothesis that explainees use a set of multimodal cues, including vocalizations, facial expressions, and movements of the torso, head, and hands, to do so. We test this hypothesis by building a random forest classifier based on a multimodal corpus of dyadic explanations (21 explainers and explainees), in which windows of understanding or non-understanding were identified by participants in a retrospective video recall task. Results show that sequences of understanding can indeed be differentiated from those of non-understanding, and that a diverse set of predictors covering a wide range of modalities contributes to this classification. Due to data sparsity and a high degree of individual variation, the generalizability of our results is currently limited, but they support our hypothesis of the relevance of multimodal display in explanatory interactions.
Olcay Türk, Stefan Lazarov, Yu Wang 0294, Hendrik Buschmeier, Angela Grimminger, Petra Wagner
ICMI6
2024 Assessing the impact of contextual framing on subjective TTS quality
abstract
Edlund J, Tånnander C, LeMaguer S, Wagner P. Assessing the impact of contextual framing on subjective TTS quality. In: Proceedings of INTERSPEECH 2024. 2024: 1205--1209.
Jens Edlund, Christina Tånnander, Sébastien Le Maguer, Petra Wagner
INTERSPEECH4
2024 Understanding "understanding": presenting a richly annotated multimodal corpus of dyadic interaction
Leonie Schade, Nico Dallmann, Olcay Türk, Stefan Lazarov, Petra Wagner
INTERSPEECH5
2023 The co-use of laughter and head gestures across speech styles
abstract
Ludusan B, Schröer M, Rossi M, Wagner P. The co-use of laughter and head gestures across speech styles. In: Interspeech 2023. Proceedings. ISCA; 2023: 3592-3596.
Bogdan Ludusan, Marin Schröer, Martina Rossi, Petra Wagner
INTERSPEECH4
2023 Effects of Meter, Genre and Experience on Pausing, Lengthening and Prosodic Phrasing in German Poetry Reading
abstract
The adequate and pleasant delivery of poetic speech remains a challenge for humans and machines alike.The present corpus study analyzes factors and strategies that characterize the stylistic expression of poetry and prose by professional actors as well as laypersons with musical training, focusing on their pausing, lengthening and intonation at verse boundaries.Our results show a clear influence on speakers' experience in modulating their speech with respect to prosodic timing: professional actors systematically insert more and more diverse prosodic boundaries and pauses than laypersons, and make strategic use of lengthening at verse endings in poetic speech.Our results further point out the relevance of pausing and lengthening as a time-buying strategy that enhances speech fluency, and we make tentative suggestions for modeling (poetic) speech in expressive speech synthesis.
Petra Wagner, Simon Betz
INTERSPEECH1
2023 The effect of conversation type on entrainment: Evidence from laughter
abstract
Entrainment is a phenomenon that occurs across several modalities and at different linguistic levels in conversation.Previous work has shown that its effects may be modulated by conversation extrinsic factors, such as the relation between the interlocutors or the speakers' traits.The current study investigates the role of conversation type on laughter entrainment.Employing dyadic interaction materials in German, containing two conversation types (free dialogues and task-based interactions), we analyzed three measures of entrainment previously proposed in the literature.The results show that the entrainment effects depend on the type of conversation, with two of the investigated measures being affected by this factor.These findings represent further evidence towards the role of situational aspects as a mediating factor in conversation.
Bogdan Ludusan, Petra Wagner
SIGDIAL2
2022 Investigation into Target Speaking Rate Adaptation for Voice Conversion
abstract
Disentangling speaker and content attributes of a speech signal into separate latent representations followed by decoding the content with an exchanged speaker representation is a popular approach for voice conversion, which can be trained with non-parallel and unlabeled speech data.However, previous approaches perform disentanglement only implicitly via some sort of information bottleneck or normalization, where it is usually hard to find a good trade-off between voice conversion and content reconstruction.Further, previous works usually do not consider an adaptation of the speaking rate to the target speaker or they put some major restrictions to the data or use case.Therefore, the contribution of this work is two-fold.First, we employ an explicit and fully unsupervised disentanglement approach, which has previously only been used for representation learning, and show that it allows to obtain both superior voice conversion and content reconstruction.Second, we investigate simple and generic approaches to linearly adapt the length of a speech signal, and hence the speaking rate, to a target speaker and show that the proposed adaptation allows to increase the speaking rate similarity with respect to the target speaker.
Michael Kuhlmann, Fritz Seebauer, Janek Ebbers, Petra Wagner, Reinhold Häb-Umbach
INTERSPEECH4
2022 Investigating phonetic convergence of laughter in conversation
abstract
Ludusan B, Schröer M, Wagner P. Investigating phonetic convergence of laughter in conversation. In: Interspeech 2022. Proceedings. ISCA: ISCA; 2022: 1332-1336.
Bogdan Ludusan, Marin Schröer, Petra Wagner
INTERSPEECH3
2022 Laughter entrainment in dyadic interactions: Temporal distribution and form
abstract
It has been established across a wide range of communicative behaviours that conversational partners tend to become more similar during their interaction. This phenomenon, often called entrainment, has been shown to take place not only at various linguistic levels, but also across different modalities. We investigated in this study whether entrainment can be found in the use of paralinguistic phenomena in conversation. Laughter is a vocalization widely recognized across cultures, and one of the most encountered paralinguistic events in spontaneous interactions. Using conversational data from three distinct languages: French, German and Mandarin Chinese, we examined two facets of entrainment: temporal and form-related. Five entrainment measures, computed across two different levels of linguistic organization, were considered in our analysis. Support was found for temporal entrainment at the laughter-token level, in how speakers of a dialogue distribute their laughter events throughout the conversation. At the turn level, speakers of all three languages showed evidence for entrainment, by aligning their laughter more with the beginning and the end of their turns. Moreover, this phenomenon seemed to be enhanced in the second half of the examined recordings, compared to the first half. The study found support also for form-related entrainment, with conversational partners employing more similar intensity levels for consecutive, than for non-consecutive laughter. Furthermore, we show that the entrainment aspects captured by our measures are independent of the degree of familiarity between the speakers.
Bogdan Ludusan, Petra Wagner
Speech Commun.2
2021 Cue Interaction in the Perception of Prosodic Prominence: The Role of Voice Quality
abstract
Ludusan B, Wagner P, Włodarczak M. Cue interaction in the perception of prosodic prominence: the role of voice quality. In: Interspeech 2021. Proceedings. ISCA; 2021: 1006-1010.
Bogdan Ludusan, Petra Wagner, Marcin Wlodarczak
Interspeech2
2021 Effects of Time Pressure and Spontaneity on Phonotactic Innovations in German Dialogues
Petra Wagner, Sina Zarrieß, Joana Cholin
Interspeech1
2020 An Evaluation of Manual and Semi-Automatic Laughter Annotation
abstract
Ludusan B, Wagner P. An Evaluation of Manual and Semi-Automatic Laughter Annotation. In: Proceedings of Interspeech 2020. ISCA; 2020: 621-625.
Bogdan Ludusan, Petra Wagner
INTERSPEECH2
2019 The Greennn Tree - Lengthening Position Influences Uncertainty Perception
abstract
Betz S, Zarrieß S, Székely É, Wagner P. The greennn tree - lengthening position influences uncertainty perception. In: Proceedings of Interspeech. 2019: 3990-3994.
Simon Betz, Sina Zarrieß, Éva Székely, Petra Wagner
INTERSPEECH4
2019 Laughter Dynamics in Dyadic Conversations
abstract
Ludusan B, Wagner P. Laughter Dynamics in Dyadic Conversations. In: Proceedings of Interspeech. 2019.
Bogdan Ludusan, Petra Wagner
INTERSPEECH2
2019 A User-Friendly and Adaptable Re-Implementation of an Acoustic Prominence Detection and Annotation Tool
Jana Voße, Petra Wagner
INTERSPEECH2
2019 Pitch Accent Trajectories Across Different Conditions of Visibility and Information Structure - Evidence from Spontaneous Dyadic Interaction
abstract
Wagner P, Bryhadyr N, Schröer M. Pitch Accent Trajectories across Different Conditions of Visibility and Information Structure - Evidence from Spontaneous Dyadic Interaction. In: Proceedings of Interspeech. ISCA; 2019: 3985-3989.
Petra Wagner, Nataliya Bryhadyr, Marin Schröer
INTERSPEECH1
2017 An adaptive neuro-fuzzy inference system for the qualitative study of perceptual prominence in linguistics
abstract
This paper explores the applications of fuzzy logic inference systems as an instrument to perform linguistic analysis in the domain of prosodic prominence. Understanding how acoustic features interact to make a linguistic unit be perceived as more relevant than the surrounding ones is generally needed to study the cognitive processes needed for speech understanding. It also has technological applications in the field of speech recognition and synthesis. We present a first experiment to show how fuzzy inference systems, being characterised by their capability to provide detailed insight about the models obtained through supervised learning can help investigate the complex relationships among acoustic features linked to prominence perception.
Autilia Vitiello, Giovanni Acampora, Francesco Cutugno, Petra Wagner, Antonio Origlia
FUZZ-IEEE4
2017 Hyperarticulation aids learning of new vowels in a developmental speech acquisition model
abstract
Many studies emphasize the importance of infant-directed speech: stronger articulated, higher-quality speech helps infants to better distinguish different speech sounds. This effect has been widely investigated in terms of the infant's perceptual capabilities, but few studies examined whether infant-directed speech has an effect on articulatory learning. In earlier studies, we developed a model that learns articulatory control for a 3D vocal tract model via goal babbling. Exploration is organized in the space of outcomes. This so called goal space is generated from a set of ambient speech sounds. Similarly to how speech from the environment shapes infant's speech perception, the data from which the goal space is learned shapes the later learning process: it determines which sounds the model is able to discriminate, and thus, which sounds it can eventually learn to produce. We investigate how speech sound quality in early learning affects the model's capability to learn new vowel sounds. The model is trained either on hyperarticulated (tense) or on hypoarticulated (lax) vowels. Then we retrain the model with vowels from the other set. Results show that new vowels can be acquired although they were not included in early learning. There is, however, an effect of learning order, showing that models first trained on the stronger articulated tense vowels easier accommodate to new vowel sounds later on.
Anja Philippsen, René Felix Reinhart, Britta Wrede, Petra Wagner
IJCNN4
2017 Increasing Recall of Lengthening Detection via Semi-Automatic Classification
abstract
Betz S, Voße J, Zarrieß S, Wagner P. Increasing Recall of Lengthening Detection via Semi-Automatic Classification. In: Proceedings of Interspeech. 2017: 1084-1088.
Simon Betz, Jana Voße, Sina Zarrieß, Petra Wagner
INTERSPEECH4
2017 What You See is What You Get Prosodically Less - Visibility Shapes Prosodic Prominence Production in Spontaneous Interaction
abstract
Wagner P, Bryhadyr N. What you see is what you get prosodically less - visibility shapes prosodic prominence production in spontaneous interaction. In: Proceedings of Interspeech 2017. 2017: 3226-3230.
Petra Wagner, Nataliya Bryhadyr
INTERSPEECH1
2016 How to Address Smart Homes with a Social Robot? A Multi-modal Corpus of User Interactions with an Intelligent Environment
Patrick Holthaus, Christian Leichsenring, Jasmin Bernotat, Viktor Richter, Marian Pohling, Birte Richter, Norman Köster, Sebastian Meyer zu Borgsen, René Zorn, Birte Schiffhauer, Kai Frederic Engelmann, Florian Lier, Simon Schulz, Philipp Cimiano, Friederike Eyssel, Thomas Hermann 0001, Franz Kummert, David Schlangen, Sven Wachsmuth, Petra Wagner, Britta Wrede, Sebastian Wrede 0001
LREC20
2015 Micro-structure of disfluencies: basics for conversational speech synthesis
abstract
Betz S, Wagner P, Schlangen D. Micro-Structure of Disfluencies: Basics for Conversational Speech Synthesis. In: Interspeech 2015. 2015: 2222-2226.
Simon Betz, Petra Wagner, David Schlangen
INTERSPEECH2
2015 Polysyllabic shortening and word-final lengthening in English
abstract
Windmann A, Simko J, Wagner P. Polysyllabic Shortening and Word-Final Lengthening in English. In: Proceedings of Interspeech 2015. 2015: 36-40.
Andreas Windmann, Juraj Simko, Petra Wagner
INTERSPEECH3
2015 Optimization-based modeling of speech timing
Andreas Windmann, Juraj Simko, Petra Wagner
Speech Commun.3
2014 A unified account of prominence effects in an optimization-based model of speech timing
abstract
Windmann A, Simko J, Wagner P. A Unified Account of Prominence Effects in an Optimization-Based Model of Speech Timing. In: Proceedings of Interspeech 2014. 2014: 159-163.
Andreas Windmann, Juraj Simko, Petra Wagner
INTERSPEECH3
2014 ALICO: a multimodal corpus for the study of active listening
Hendrik Buschmeier, Zofia Malisz, Joanna Skubisz, Marcin Wlodarczak, Ipke Wachsmuth, Stefan Kopp, Petra Wagner
LREC7
2014 Gesture and speech in interaction: An overview
Petra Wagner, Zofia Malisz, Stefan Kopp
Speech Commun.1
2013 Timing and entrainment of multimodal backchanneling behavior for an embodied conversational agent
abstract
We report on an analysis of feedback behavior in an Active Listening Corpus as produced verbally, visually (head movement) and bimodally. The behavior is modeled in an embodied conversational agent and displayed in a conversation with a real human to human participants for perceptual evaluation. Five strategies for the timing of backchannels are compared: copying the timing of the original human listener, producing backchannels at randomly selected times, producing backchannels according to high level timing distributions relative to the interlocutor's utterance and pauses, or according to local entrainment to the interlocutors' vowels, or according to both. Human observers judge that models with global timing distributions miss less opportunities for backchanneling than random timing.
Benjamin Inden, Zofia Malisz, Petra Wagner, Ipke Wachsmuth
ICMI3
2013 Using generalized additive models and random forests to model prosodic prominence in German
abstract
The perception of prosodic prominence is influenced by different sources like different acoustic cues, linguistic expectations and context. We use a generalized additive model and a random forest to model the perceived prominence on a corpus of spoken German. Both models are able to explain over 80% of the variance. While the random forests give us some insights on the relative importance of the cues, the general additive model gives us insights on the interaction between different cues to prominence.
Denis Arnold, Petra Wagner, R. Harald Baayen
INTERSPEECH2
2013 Effects of lexical class and lemma frequency on German homographs
abstract
German demonstrative pronouns, relative pronouns, and definite articles are segmentally identical but differ strongly in the frequency with which they appear. We examined the production of five such particles in a reading task. In a comparison of orthographically identical word pairs belonging to different lexical classes we found small but significant differences in word and vowel duration, prominence, and spectral similarity. Three of the particles in particular tended to be longer and more prominent when they occurred as demonstrative or relative articles than when they were assigned their usual role as definite articles. Index Terms: lemma frequency, duration, prominence 1.
Barbara Samlowski, Petra Wagner, Bernd Möbius
INTERSPEECH2
2013 Modeling durational incompressibility
abstract
We show how incompressibility, a well-described property of some prosodic timing effects, can be accounted for in an optimization-based model of speech timing.Preliminary results of a corpus study are presented, replicating and generalizing previous findings on incompressibility as a function of increasing speaking rate.We then introduce the architecture of our model and present results of simulation experiments that reproduce the results of the corpus analysis.Results suggest that incompressibility can be interpreted as a consequence of tradeoffs between competing requirements of production efficiency and communicative efficacy.
Andreas Windmann, Juraj Simko, Britta Wrede, Petra Wagner
INTERSPEECH4
2013 Pitch and duration as a basis for entrainment of overlapped speech onsets
abstract
The present paper reports on the impact of pitch accents and duration on temporal organisation of overlapping speech onsets in spontaneous dialogue.We observe a non-random pattern of overlap initiations within intervals between consecutive pitch accents, thus extending our earlier reports of a similar effect within vowel-to-vowel intervals.The latter finding was interpreted as a tendency to start overlapped speech directly before perceptually prominent vocalic onsets.In an attempt to reconcile these results, we investigate whether the effect observed on vowel-to-vowel intervals is influenced by presence of pitch accents and lengthening, both of which are known to be correlated with perceptual prominence.We find a strong effect of duration, which, however, does not on its own account fully for the observed pattern, indicating that other correlates of prominence might be involved in guiding the timing of overlap onsets.
Marcin Wlodarczak, Juraj Simko, Petra Wagner
INTERSPEECH3
2013 Effects of talk-spurt silence boundary thresholds on distribution of gaps and overlaps
abstract
Faced with lack of objective and easily applicable criteria for segmentation of speech into dialogue turns, many authors resort instead to units defined in terms of stretches of speech minimally bounded by silence of some predefined duration. There is, however, no consensus concerning silence thresholds employed. While such thresholds can be established on perceptual grounds, in practice a wide range of values is used. As this has a direct impact on the reported frequencies of silences and overlaps, the discrepancies make comparisons of results across different studies difficult. In an attempt to overcome these problems in the present paper we use the Switchboard corpus to evaluate the expected variability in distributions of inter- and intra-speaker intervals when silence boundary thresholds of inter-pausal units are manipulated. Index Terms: dialogue segmentation, inter-pausal units, gaps and overlaps
Marcin Wlodarczak, Petra Wagner
INTERSPEECH2
2012 Rapid entrainment to spontaneous speech: A comparison of oscillator models
Benjamin Inden, Zofia Malisz, Petra Wagner, Ipke Wachsmuth
CogSci3
2012 Obtaining prominence judgments from naïve listeners - Influence of rating scales, linguistic levels and normalisation
abstract
A frequently replicated finding is that higher frequency words tend to be shorter and contain more strongly reduced vowels. However, little is known about potential differences in the articulatory gestures for high vs. low frequency words. The present study made use of electromagnetic articulography to investigate the production of two German vowels, [i] and [a], embedded in high and low frequency words. We found that word frequency differently affected the production of [i] and [a] at the temporal as well as the gestural level. Higher frequency of use predicted greater acoustic durations for long vowels; reduced durations for short vowels; articulatory trajectories with greater tongue height for [i] and more pronounced downward articulatory trajectories for [a]. These results show that the phonological contrast between short and long vowels is learned better with experience, and challenge both the Smooth Signal Redundancy Hypothesis and current theories of German phonology.
Denis Arnold, Petra Wagner, Bernd Möbius
INTERSPEECH2
2012 Gaze Patterns in Turn-Taking
abstract
Oertel C, Wlodarczak M, Edlund J, Wagner P, Gustafson J. Gaze patterns in turn-taking. In: 13th Annual Conference of the International Speech Communication Association 2012 (INTERSPEECH 2012). Red Hook, NY: Curran; 2013: 2243-2246.
Catharine Oertel, Marcin Wlodarczak, Jens Edlund, Petra Wagner, Joakim Gustafson
INTERSPEECH4
2012 Disentangling lexical, morphological, syntactic and semantic influences on German prominence - Evidence from a production study
abstract
Samlowski B, Wagner P, Möbius B. Disentangling lexical, morphological, syntactic and semantic influences on German prominence – Evidence from a production study. In: Proceedings of Interspeech 2012. 2012: 2406-2409.
Barbara Samlowski, Petra Wagner, Bernd Möbius
INTERSPEECH2
2012 Objective, Subjective and Linguistic Roads to Perceptual Prominence - How are they compared and why?
abstract
Prosodic prominence denotes the perceptual salience of linguistic units. There exists no agreement on (1) ade-quate methods for its subjective measurement, (2) its ob-jective acoustic correlates and (3) its relationship to lin-guistic structure. A traditional approach for evaluating any of these descriptive layers is an inter-level compari-son, e.g. between a perceptual and an acoustic model of prominence. However, (1) there exists no standard pro-cedure for such a comparison, and (2) such a comparison is misleading if both layers are expected to be symmetri-cal, given the neglected influence of linguistic top-down expectancies. We propose an evaluation procedure for prominence models relying on tripartite correlations of perception, its acoustic correlates and linguistic expecta-tions. We suggest a novel correlation metric and test its usefulness on a prosodic corpus of German. Index Terms: prominence, evaluation, prosody 1. Introduction: Model
Petra Wagner, Fabio Tamburini, Andreas Windmann
INTERSPEECH1
2012 Temporal entrainment in overlapped speech: Cross-linguistic study
abstract
Wlodarczak M, Simko J, Wagner P. Temporal entrainment in overlapped speech: Cross-linguistic study. In: 13th Annual Conference of the International Speech Communication Association 2012 (INTERSPEECH 2012). Vol. 1. Red Hook, NY: Curran; 2013: 614-617.
Marcin Wlodarczak, Juraj Simko, Petra Wagner
INTERSPEECH3
2011 Dynamic perception-production oscillation model in human-machine communication
abstract
The goal of the present article is to introduce a new concept of a perception-production timing model in human-machine communication. The model implements a low-level cognitive timing and coordination mechanism. The basic element of the model is a dynamic oscillator capable of tracking reoccurring events in time. The organization of the oscillators in a network is being referred to as the Dynamic Perception-Production Oscillation Model (DPPOM). The DPPOM is largely based on findings in psychological and phonetic experiments on timing in speech perception and production. It consists of two sub-systems, a perception sub-system and a production sub-system. The perception sub-system accounts for information clustering in an input sequence of events. The production sub-system accounts for speech production rhythmically entrained to the input sequence. We propose a system architecture integrating both sub-systems, providing a flexible mechanism for perception-production timing in dialogues. The model's functionality was evaluated in two experiments.
Igor Jauk, Ipke Wachsmuth, Petra Wagner
ICMI3
2011 Comparing Word and Syllable Prominence Rated by Naïve Listeners
abstract
Arnold D, Möbius B, Wagner P. Comparing word and syllable prominence rated by naive listeners. In: Proceedings of Interspeech 2011. 2011: 1877-1880.
Denis Arnold, Bernd Möbius, Petra Wagner
INTERSPEECH3
2011 'Are You Sure You're Paying Attention?' - 'Uh-Huh' Communicating Understanding as a Marker of Attentiveness
abstract
Buschmeier H, Malisz Z, Wlodarczak M, Kopp S, Wagner P. 'Are you sure you're paying attention?' – 'Uh-huh'. Communicating understanding as a marker of attentiveness. In: Proceedings of INTERSPEECH 2011. International Speech Communication Association; 2011: 2057-2060.
Hendrik Buschmeier, Zofia Malisz, Marcin Wlodarczak, Stefan Kopp, Petra Wagner
INTERSPEECH5
2011 Comparing Syllable Frequencies in Corpora of Written and Spoken Language
abstract
In the study, various German language corpora were compared in order to discover the extent to which syllable frequencies remain stable across different contexts and modalities. Although considerable differences in relative frequency were found among the more common syllables, rank numbers proved to be more robust. Variation across corpora was mostly due to vocabulary characteristics of particular corpus domains rather than to systematic differences between spoken and written language. The results indicate that syllable frequencies in written corpora can be taken as a rough estimate for their frequency in spoken language. Index Terms: syllabary, syllable frequencies, spoken and written language corpora
Barbara Samlowski, Bernd Möbius, Petra Wagner
INTERSPEECH3
2011 Using Prominence Detection to Generate Acoustic Feedback in Tutoring Scenarios
abstract
Robots interacting with humans need to understand actions and make use of language in social interactions. Research on infant development has shown that language helps the learner to structure visual observations of action. This acoustic information typically in the form of narration overlaps with action sequences and provides infants with a bottom-up guide to find structure within them. This concept has been introduced as acoustic packaging by Hirsh-Pasek and Golinkoff. We developed and integrated a prominence detection module in our acoustic packaging system to detect semantically relevant information linguistically\nhighlighted by the tutor. Evaluation results on speech data from adult-infant interactions show a significant agreement with human raters. Furthermore a first approach based on acoustic packages which uses the prominence detection results to generate acoustic feedback is presented.\n\nIndex Terms: prominence, multimodal action segmentation,\nhuman robot interaction, feedback
Lars Schillingmann, Petra Wagner, Christian Munier, Britta Wrede, Katharina J. Rohlfing
INTERSPEECH2
2011 Prominence-Based Prosody Prediction for Unit Selection Speech Synthesis
abstract
This paper describes the development and evaluation of a\nprosody prediction module for unit selection speech synthesis that is based on the notion of perceptual prominence. We outline the design principles of the module and describe its implementation in the Bonn Open Synthesis System (BOSS).\nMoreover, we report results of perception experiments that\nhave been conducted in order to evaluate prominence\nprediction. The paper is concluded by a general discussion of the approach and a sketch of perspectives for further work.
Andreas Windmann, Igor Jauk, Fabio Tamburini, Petra Wagner
INTERSPEECH4
2009 Assessing a speaker for fast speech in unit selection speech synthesis
abstract
Moers D, Wagner P. Assessing a speaker for fast speech in unit selection speech synthesis. In: Proceedings of Interspeech. 2009: 2071-2074.
Donata Moers, Petra Wagner
INTERSPEECH2
2009 Paper 8003 was not available at the time of publication oral presentation of poster papers no time to lose? time shrinking effects enhance the impression of rhythmic "isochrony" and fast speech rate
Petra Wagner, Andreas Windmann
INTERSPEECH1
2007 On automatic prominence detection for German
abstract
Perceptual prominence is an important indicator of a word's and syllable's lexical, syntactic, \nsemantic and pragmatic status in a discourse. Its automatic annotation would be a valuable \nenrichment of large databases used in unit selection speech synthesis and speech recognition. \nWhile much research has been carried out on the interaction between prominence and \nacoustic factors, little progress has been made in its automatic annotation. Previous \napproaches to German relied on linguistic features in prominence detection, but a purely \nacoustic method would be advantageous. We applied an algorithm to German data that had \nbeen previously used for English and Italian. Both the algorithm and the data annotation \nencode prominence as a continuous rather than a categorical parameter. First results are \nencouraging, but again show that prominence perception relies on linguistic expectancies as \nwell as acoustic patterns. Also, our results further strengthen the view that force accents are a \nmore reliable cue to prominence than pitch accents in German.
Fabio Tamburini, Petra Wagner
INTERSPEECH2
2005 Great expectations - introspective vs. perceptual prominence ratings and their acoustic correlates
abstract
In order to gain knowledge about the interaction between top-down expectations of listeners concerning prosodic prominence and its acoustic correlates, two exploratory empirical studies were carried out. First, native and non-native subjects rated prominences of speech read at normal and very fast - prosodically very different - speech. Later, these ratings were compared with introspective prominence ratings of different listeners. First results indicate a major influence of the introspection on prominence ratings, especially if acoustic cues are difficult to interpret, as it is the case in very fast speech. Compared to native subjects, non-natives rely less on their introspection and more on the acoustics.
Petra Wagner
INTERSPEECH1
2004 Bonntempo-corpus and bonntempo-tools: a database for the study of speech rhythm and rate
abstract
Work is currently being carried out on a speech database constructed in order to study speech rhythm in connection with speech rate. The database, BonnTempo-Corpus, and the Praat based analysis tools, BonnTempo-Tools, are a powerful instrument for examining various aspects of recently proposed rhythm measures (e.g. %V, C, nPVI, rPVI, etc.) in relation to speech rate among a wide range of languages and speakers. First observations pose new problems on traditionally not well classifiable languages like Czech.
Volker Dellwo, Bianca Aschenberner, Petra Wagner, Jana Dancovicova, Ingmar Steiner
INTERSPEECH3
2001 Speech synthesis development made easy: the bonn open synthesis system
abstract
This paper describes a new open source architecture for unit-selection based speech synthesis called BOSS (Bonn Open Synthesis System). It is built up modularly, with communications between modules taking place in a fixed format. This makes the addition, deletion and substitution of modules very easy. The strict separation between data and algorithms allows for the simple creation of new speech corpora for different domains and languages. 1.
Esther Klabbers, Karlheinz Stöber, Raymond N. J. Veldhuis, Petra Wagner, Stefan Breuer
INTERSPEECH4
1999 Synthesis by word concatenation
abstract
Stöber K, Portele T, Wagner P, Hess W. Synthesis by Word Concatenation. In: Proceedings of Interspeech 1999. Vol 2. Budapest, Hungary; 1999: 619-622.
Karlheinz Stöber, Thomas Portele, Petra Wagner, Wolfgang Hess
EUROSPEECH3