EDBT 2026 Demo / reviewers in the wild / expert
Petra Wagner
dblp:42/9238
· DBLP profile ↗
56ranked-venue papers
8as first author
15since 2021 · last 2025
0000-0001-6662-3612ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 47 · 8 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speech Synthesis along Perceptual Voice Quality DimensionsabstractWhile expressive speech synthesis or voice conversion systems mainly focus on controlling or manipulating abstract prosodic characteristics of speech, such as emotion or accent, we here address the control of perceptual voice qualities (PVQs) recognized by phonetic experts, which are speech properties at a lower level of abstraction. The ability to manipulate PVQs can be a valuable tool for teaching speech pathologists in training or voice actors. In this paper, we integrate a Conditional Continuous-Normalizing-Flow-based method into a Text-to-Speech system to modify perceptual voice attributes on a continuous scale. Unlike previous approaches, our system avoids direct manipulation of acoustic correlates and instead learns from examples. We demonstrate the system's capability by manipulating four voice qualities: Roughness, breathiness, resonance and weight. Phonetic experts evaluated these modifications, both for seen and unseen speaker conditions. The results highlight both the system's strengths and areas for improvement. Frederik Rautenberg, Michael Kuhlmann, Fritz Seebauer, Jana Wiechmann, Petra Wagner, Reinhold Häb-Umbach |
ICASSP | 5 |
| 2025 | Towards Frame-level Quality Predictions of Synthetic SpeechabstractKuhlmann M, Seebauer FM, Wagner P, Haeb-Umbach R. Towards Frame-level Quality Predictions of Synthetic Speech. In: Interspeech 2025. Interspeech. Baixas: International Speech Communication Association; 2025: 2300-2304. Michael Kuhlmann, Fritz Seebauer, Petra Wagner, Reinhold Häb-Umbach |
INTERSPEECH | 3 |
| 2025 | Synthesizing Speech with Selected Perceptual Voice Qualities - A Case Study with Creaky VoiceabstractRautenberg F, Seebauer FM, Wiechmann J, Kuhlmann M, Wagner P, Haeb-Umbach R. Synthesizing Speech with Selected Perceptual Voice Qualities – A Case Study with Creaky Voice. In: Interspeech 2025. ISCA: ISCA; 2025: 1633-1637. Frederik Rautenberg, Fritz Seebauer, Jana Wiechmann, Michael Kuhlmann, Petra Wagner, Reinhold Häb-Umbach |
INTERSPEECH | 5 |
| 2025 | A comparison of discrete and continuous prominence perception methods in GermanabstractIn this paper we report on three methods to investigate syllable-based prominence identification on a set of German read sentences: prosodic expert annotations of pitch accentuation, as well as a Rapid Prosody Transcription (RPT) style task and a finger-tapping task performed by naive listeners. In the present study, audio recordings of the speech materials used to elicit prominence judgments were supplemented by signals collected with miniature accelerometers placed on the throat skin below the glottis , allowing for a more reliable investigation of the contribution of voice quality. Various other signal-based parameters are correlated with prominence judgments to confirm if findings from previous work on word-based prominence judgments also hold for judgments at the syllable level. Results replicate past findings for German and other languages: the presence and type of pitch accentuation are reliable predictors of prominence judgments by naive listeners, on both the RPT and tapping task. Various individual acoustic parameters such as f0, duration and intensity were once again found to systematically covary with greater perceived prominence. A direct comparison of tapping and RPT results moreover indicated that beyond pitch accent related factors, listeners may employ different strategies to judge prominence as a function of task. In the RPT, they rely more on their knowledge of whether a given syllable carries lexical stress, whereas in the tapping task, they attend relatively strongly to acoustic duration. It is also shown that voice quality varies along with prominence ratings, but less strongly than other features such as duration. Anna Bruggeman, Marcin Wlodarczak, Petra Wagner |
Speech Commun. | 3 |
| 2024 | Predictability of Understanding in Explanatory Interactions Based on Multimodal CuesabstractIn explanatory interactions, explainees are expected to continuously provide feedback to explainers by signaling whether they understand an ongoing explanation. The study presented in this paper is based on the hypothesis that explainees use a set of multimodal cues, including vocalizations, facial expressions, and movements of the torso, head, and hands, to do so. We test this hypothesis by building a random forest classifier based on a multimodal corpus of dyadic explanations (21 explainers and explainees), in which windows of understanding or non-understanding were identified by participants in a retrospective video recall task. Results show that sequences of understanding can indeed be differentiated from those of non-understanding, and that a diverse set of predictors covering a wide range of modalities contributes to this classification. Due to data sparsity and a high degree of individual variation, the generalizability of our results is currently limited, but they support our hypothesis of the relevance of multimodal display in explanatory interactions. Olcay Türk, Stefan Lazarov, Yu Wang 0294, Hendrik Buschmeier, Angela Grimminger, Petra Wagner |
ICMI | 6 |
| 2024 | Assessing the impact of contextual framing on subjective TTS qualityabstractEdlund J, Tånnander C, LeMaguer S, Wagner P. Assessing the impact of contextual framing on subjective TTS quality. In: Proceedings of INTERSPEECH 2024. 2024: 1205--1209. Jens Edlund, Christina Tånnander, Sébastien Le Maguer, Petra Wagner |
INTERSPEECH | 4 |
| 2024 | Understanding "understanding": presenting a richly annotated multimodal corpus of dyadic interaction
Leonie Schade, Nico Dallmann, Olcay Türk, Stefan Lazarov, Petra Wagner |
INTERSPEECH | 5 |
| 2023 | The co-use of laughter and head gestures across speech stylesabstractLudusan B, Schröer M, Rossi M, Wagner P. The co-use of laughter and head gestures across speech styles. In: Interspeech 2023. Proceedings. ISCA; 2023: 3592-3596. Bogdan Ludusan, Marin Schröer, Martina Rossi, Petra Wagner |
INTERSPEECH | 4 |
| 2023 | Effects of Meter, Genre and Experience on Pausing, Lengthening and Prosodic Phrasing in German Poetry ReadingabstractThe adequate and pleasant delivery of poetic speech remains a challenge for humans and machines alike.The present corpus study analyzes factors and strategies that characterize the stylistic expression of poetry and prose by professional actors as well as laypersons with musical training, focusing on their pausing, lengthening and intonation at verse boundaries.Our results show a clear influence on speakers' experience in modulating their speech with respect to prosodic timing: professional actors systematically insert more and more diverse prosodic boundaries and pauses than laypersons, and make strategic use of lengthening at verse endings in poetic speech.Our results further point out the relevance of pausing and lengthening as a time-buying strategy that enhances speech fluency, and we make tentative suggestions for modeling (poetic) speech in expressive speech synthesis. Petra Wagner, Simon Betz |
INTERSPEECH | 1 |
| 2023 | The effect of conversation type on entrainment: Evidence from laughterabstractEntrainment is a phenomenon that occurs across several modalities and at different linguistic levels in conversation.Previous work has shown that its effects may be modulated by conversation extrinsic factors, such as the relation between the interlocutors or the speakers' traits.The current study investigates the role of conversation type on laughter entrainment.Employing dyadic interaction materials in German, containing two conversation types (free dialogues and task-based interactions), we analyzed three measures of entrainment previously proposed in the literature.The results show that the entrainment effects depend on the type of conversation, with two of the investigated measures being affected by this factor.These findings represent further evidence towards the role of situational aspects as a mediating factor in conversation. Bogdan Ludusan, Petra Wagner |
SIGDIAL | 2 |
| 2022 | Investigation into Target Speaking Rate Adaptation for Voice ConversionabstractDisentangling speaker and content attributes of a speech signal into separate latent representations followed by decoding the content with an exchanged speaker representation is a popular approach for voice conversion, which can be trained with non-parallel and unlabeled speech data.However, previous approaches perform disentanglement only implicitly via some sort of information bottleneck or normalization, where it is usually hard to find a good trade-off between voice conversion and content reconstruction.Further, previous works usually do not consider an adaptation of the speaking rate to the target speaker or they put some major restrictions to the data or use case.Therefore, the contribution of this work is two-fold.First, we employ an explicit and fully unsupervised disentanglement approach, which has previously only been used for representation learning, and show that it allows to obtain both superior voice conversion and content reconstruction.Second, we investigate simple and generic approaches to linearly adapt the length of a speech signal, and hence the speaking rate, to a target speaker and show that the proposed adaptation allows to increase the speaking rate similarity with respect to the target speaker. Michael Kuhlmann, Fritz Seebauer, Janek Ebbers, Petra Wagner, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2022 | Investigating phonetic convergence of laughter in conversationabstractLudusan B, Schröer M, Wagner P. Investigating phonetic convergence of laughter in conversation. In: Interspeech 2022. Proceedings. ISCA: ISCA; 2022: 1332-1336. Bogdan Ludusan, Marin Schröer, Petra Wagner |
INTERSPEECH | 3 |
| 2022 | Laughter entrainment in dyadic interactions: Temporal distribution and formabstractIt has been established across a wide range of communicative behaviours that conversational partners tend to become more similar during their interaction. This phenomenon, often called entrainment, has been shown to take place not only at various linguistic levels, but also across different modalities. We investigated in this study whether entrainment can be found in the use of paralinguistic phenomena in conversation. Laughter is a vocalization widely recognized across cultures, and one of the most encountered paralinguistic events in spontaneous interactions. Using conversational data from three distinct languages: French, German and Mandarin Chinese, we examined two facets of entrainment: temporal and form-related. Five entrainment measures, computed across two different levels of linguistic organization, were considered in our analysis. Support was found for temporal entrainment at the laughter-token level, in how speakers of a dialogue distribute their laughter events throughout the conversation. At the turn level, speakers of all three languages showed evidence for entrainment, by aligning their laughter more with the beginning and the end of their turns. Moreover, this phenomenon seemed to be enhanced in the second half of the examined recordings, compared to the first half. The study found support also for form-related entrainment, with conversational partners employing more similar intensity levels for consecutive, than for non-consecutive laughter. Furthermore, we show that the entrainment aspects captured by our measures are independent of the degree of familiarity between the speakers. Bogdan Ludusan, Petra Wagner |
Speech Commun. | 2 |
| 2021 | Cue Interaction in the Perception of Prosodic Prominence: The Role of Voice QualityabstractLudusan B, Wagner P, Włodarczak M. Cue interaction in the perception of prosodic prominence: the role of voice quality. In: Interspeech 2021. Proceedings. ISCA; 2021: 1006-1010. Bogdan Ludusan, Petra Wagner, Marcin Wlodarczak |
Interspeech | 2 |
| 2021 | Effects of Time Pressure and Spontaneity on Phonotactic Innovations in German Dialogues
Petra Wagner, Sina Zarrieß, Joana Cholin |
Interspeech | 1 |
| 2020 | An Evaluation of Manual and Semi-Automatic Laughter AnnotationabstractLudusan B, Wagner P. An Evaluation of Manual and Semi-Automatic Laughter Annotation. In: Proceedings of Interspeech 2020. ISCA; 2020: 621-625. Bogdan Ludusan, Petra Wagner |
INTERSPEECH | 2 |
| 2019 | The Greennn Tree - Lengthening Position Influences Uncertainty PerceptionabstractBetz S, Zarrieß S, Székely É, Wagner P. The greennn tree - lengthening position influences uncertainty perception. In: Proceedings of Interspeech. 2019: 3990-3994. Simon Betz, Sina Zarrieß, Éva Székely, Petra Wagner |
INTERSPEECH | 4 |
| 2019 | Laughter Dynamics in Dyadic ConversationsabstractLudusan B, Wagner P. Laughter Dynamics in Dyadic Conversations. In: Proceedings of Interspeech. 2019. Bogdan Ludusan, Petra Wagner |
INTERSPEECH | 2 |
| 2019 | A User-Friendly and Adaptable Re-Implementation of an Acoustic Prominence Detection and Annotation Tool
Jana Voße, Petra Wagner |
INTERSPEECH | 2 |
| 2019 | Pitch Accent Trajectories Across Different Conditions of Visibility and Information Structure - Evidence from Spontaneous Dyadic InteractionabstractWagner P, Bryhadyr N, Schröer M. Pitch Accent Trajectories across Different Conditions of Visibility and Information Structure - Evidence from Spontaneous Dyadic Interaction. In: Proceedings of Interspeech. ISCA; 2019: 3985-3989. Petra Wagner, Nataliya Bryhadyr, Marin Schröer |
INTERSPEECH | 1 |
| 2017 | An adaptive neuro-fuzzy inference system for the qualitative study of perceptual prominence in linguisticsabstractThis paper explores the applications of fuzzy logic inference systems as an instrument to perform linguistic analysis in the domain of prosodic prominence. Understanding how acoustic features interact to make a linguistic unit be perceived as more relevant than the surrounding ones is generally needed to study the cognitive processes needed for speech understanding. It also has technological applications in the field of speech recognition and synthesis. We present a first experiment to show how fuzzy inference systems, being characterised by their capability to provide detailed insight about the models obtained through supervised learning can help investigate the complex relationships among acoustic features linked to prominence perception. Autilia Vitiello, Giovanni Acampora, Francesco Cutugno, Petra Wagner, Antonio Origlia |
FUZZ-IEEE | 4 |
| 2017 | Hyperarticulation aids learning of new vowels in a developmental speech acquisition modelabstractMany studies emphasize the importance of infant-directed speech: stronger articulated, higher-quality speech helps infants to better distinguish different speech sounds. This effect has been widely investigated in terms of the infant's perceptual capabilities, but few studies examined whether infant-directed speech has an effect on articulatory learning. In earlier studies, we developed a model that learns articulatory control for a 3D vocal tract model via goal babbling. Exploration is organized in the space of outcomes. This so called goal space is generated from a set of ambient speech sounds. Similarly to how speech from the environment shapes infant's speech perception, the data from which the goal space is learned shapes the later learning process: it determines which sounds the model is able to discriminate, and thus, which sounds it can eventually learn to produce. We investigate how speech sound quality in early learning affects the model's capability to learn new vowel sounds. The model is trained either on hyperarticulated (tense) or on hypoarticulated (lax) vowels. Then we retrain the model with vowels from the other set. Results show that new vowels can be acquired although they were not included in early learning. There is, however, an effect of learning order, showing that models first trained on the stronger articulated tense vowels easier accommodate to new vowel sounds later on. Anja Philippsen, René Felix Reinhart, Britta Wrede, Petra Wagner |
IJCNN | 4 |
| 2017 | Increasing Recall of Lengthening Detection via Semi-Automatic ClassificationabstractBetz S, Voße J, Zarrieß S, Wagner P. Increasing Recall of Lengthening Detection via Semi-Automatic Classification. In: Proceedings of Interspeech. 2017: 1084-1088. Simon Betz, Jana Voße, Sina Zarrieß, Petra Wagner |
INTERSPEECH | 4 |
| 2017 | What You See is What You Get Prosodically Less - Visibility Shapes Prosodic Prominence Production in Spontaneous InteractionabstractWagner P, Bryhadyr N. What you see is what you get prosodically less - visibility shapes prosodic prominence production in spontaneous interaction. In: Proceedings of Interspeech 2017. 2017: 3226-3230. Petra Wagner, Nataliya Bryhadyr |
INTERSPEECH | 1 |
| 2016 | How to Address Smart Homes with a Social Robot? A Multi-modal Corpus of User Interactions with an Intelligent Environment
Patrick Holthaus, Christian Leichsenring, Jasmin Bernotat, Viktor Richter, Marian Pohling, Birte Richter, Norman Köster, Sebastian Meyer zu Borgsen, René Zorn, Birte Schiffhauer, Kai Frederic Engelmann, Florian Lier, Simon Schulz, Philipp Cimiano, Friederike Eyssel, Thomas Hermann 0001, Franz Kummert, David Schlangen, Sven Wachsmuth, Petra Wagner, Britta Wrede, Sebastian Wrede 0001 |
LREC | 20 |
| 2015 | Micro-structure of disfluencies: basics for conversational speech synthesisabstractBetz S, Wagner P, Schlangen D. Micro-Structure of Disfluencies: Basics for Conversational Speech Synthesis. In: Interspeech 2015. 2015: 2222-2226. Simon Betz, Petra Wagner, David Schlangen |
INTERSPEECH | 2 |
| 2015 | Polysyllabic shortening and word-final lengthening in EnglishabstractWindmann A, Simko J, Wagner P. Polysyllabic Shortening and Word-Final Lengthening in English. In: Proceedings of Interspeech 2015. 2015: 36-40. Andreas Windmann, Juraj Simko, Petra Wagner |
INTERSPEECH | 3 |
| 2015 | Optimization-based modeling of speech timing
Andreas Windmann, Juraj Simko, Petra Wagner |
Speech Commun. | 3 |
| 2014 | A unified account of prominence effects in an optimization-based model of speech timingabstractWindmann A, Simko J, Wagner P. A Unified Account of Prominence Effects in an Optimization-Based Model of Speech Timing. In: Proceedings of Interspeech 2014. 2014: 159-163. Andreas Windmann, Juraj Simko, Petra Wagner |
INTERSPEECH | 3 |
| 2014 | ALICO: a multimodal corpus for the study of active listening
Hendrik Buschmeier, Zofia Malisz, Joanna Skubisz, Marcin Wlodarczak, Ipke Wachsmuth, Stefan Kopp, Petra Wagner |
LREC | 7 |
| 2014 | Gesture and speech in interaction: An overview
Petra Wagner, Zofia Malisz, Stefan Kopp |
Speech Commun. | 1 |
| 2013 | Timing and entrainment of multimodal backchanneling behavior for an embodied conversational agentabstractWe report on an analysis of feedback behavior in an Active Listening Corpus as produced verbally, visually (head movement) and bimodally. The behavior is modeled in an embodied conversational agent and displayed in a conversation with a real human to human participants for perceptual evaluation. Five strategies for the timing of backchannels are compared: copying the timing of the original human listener, producing backchannels at randomly selected times, producing backchannels according to high level timing distributions relative to the interlocutor's utterance and pauses, or according to local entrainment to the interlocutors' vowels, or according to both. Human observers judge that models with global timing distributions miss less opportunities for backchanneling than random timing. Benjamin Inden, Zofia Malisz, Petra Wagner, Ipke Wachsmuth |
ICMI | 3 |
| 2013 | Using generalized additive models and random forests to model prosodic prominence in GermanabstractThe perception of prosodic prominence is influenced by different sources like different acoustic cues, linguistic expectations and context. We use a generalized additive model and a random forest to model the perceived prominence on a corpus of spoken German. Both models are able to explain over 80% of the variance. While the random forests give us some insights on the relative importance of the cues, the general additive model gives us insights on the interaction between different cues to prominence. Denis Arnold, Petra Wagner, R. Harald Baayen |
INTERSPEECH | 2 |
| 2013 | Effects of lexical class and lemma frequency on German homographsabstractGerman demonstrative pronouns, relative pronouns, and definite articles are segmentally identical but differ strongly in the frequency with which they appear. We examined the production of five such particles in a reading task. In a comparison of orthographically identical word pairs belonging to different lexical classes we found small but significant differences in word and vowel duration, prominence, and spectral similarity. Three of the particles in particular tended to be longer and more prominent when they occurred as demonstrative or relative articles than when they were assigned their usual role as definite articles. Index Terms: lemma frequency, duration, prominence 1. Barbara Samlowski, Petra Wagner, Bernd Möbius |
INTERSPEECH | 2 |
| 2013 | Modeling durational incompressibilityabstractWe show how incompressibility, a well-described property of some prosodic timing effects, can be accounted for in an optimization-based model of speech timing.Preliminary results of a corpus study are presented, replicating and generalizing previous findings on incompressibility as a function of increasing speaking rate.We then introduce the architecture of our model and present results of simulation experiments that reproduce the results of the corpus analysis.Results suggest that incompressibility can be interpreted as a consequence of tradeoffs between competing requirements of production efficiency and communicative efficacy. Andreas Windmann, Juraj Simko, Britta Wrede, Petra Wagner |
INTERSPEECH | 4 |
| 2013 | Pitch and duration as a basis for entrainment of overlapped speech onsetsabstractThe present paper reports on the impact of pitch accents and duration on temporal organisation of overlapping speech onsets in spontaneous dialogue.We observe a non-random pattern of overlap initiations within intervals between consecutive pitch accents, thus extending our earlier reports of a similar effect within vowel-to-vowel intervals.The latter finding was interpreted as a tendency to start overlapped speech directly before perceptually prominent vocalic onsets.In an attempt to reconcile these results, we investigate whether the effect observed on vowel-to-vowel intervals is influenced by presence of pitch accents and lengthening, both of which are known to be correlated with perceptual prominence.We find a strong effect of duration, which, however, does not on its own account fully for the observed pattern, indicating that other correlates of prominence might be involved in guiding the timing of overlap onsets. Marcin Wlodarczak, Juraj Simko, Petra Wagner |
INTERSPEECH | 3 |
| 2013 | Effects of talk-spurt silence boundary thresholds on distribution of gaps and overlapsabstractFaced with lack of objective and easily applicable criteria for segmentation of speech into dialogue turns, many authors resort instead to units defined in terms of stretches of speech minimally bounded by silence of some predefined duration. There is, however, no consensus concerning silence thresholds employed. While such thresholds can be established on perceptual grounds, in practice a wide range of values is used. As this has a direct impact on the reported frequencies of silences and overlaps, the discrepancies make comparisons of results across different studies difficult. In an attempt to overcome these problems in the present paper we use the Switchboard corpus to evaluate the expected variability in distributions of inter- and intra-speaker intervals when silence boundary thresholds of inter-pausal units are manipulated. Index Terms: dialogue segmentation, inter-pausal units, gaps and overlaps Marcin Wlodarczak, Petra Wagner |
INTERSPEECH | 2 |
| 2012 | Rapid entrainment to spontaneous speech: A comparison of oscillator models
Benjamin Inden, Zofia Malisz, Petra Wagner, Ipke Wachsmuth |
CogSci | 3 |
| 2012 | Obtaining prominence judgments from naïve listeners - Influence of rating scales, linguistic levels and normalisationabstractA frequently replicated finding is that higher frequency words tend to be shorter and contain more strongly reduced vowels. However, little is known about potential differences in the articulatory gestures for high vs. low frequency words. The present study made use of electromagnetic articulography to investigate the production of two German vowels, [i] and [a], embedded in high and low frequency words. We found that word frequency differently affected the production of [i] and [a] at the temporal as well as the gestural level. Higher frequency of use predicted greater acoustic durations for long vowels; reduced durations for short vowels; articulatory trajectories with greater tongue height for [i] and more pronounced downward articulatory trajectories for [a]. These results show that the phonological contrast between short and long vowels is learned better with experience, and challenge both the Smooth Signal Redundancy Hypothesis and current theories of German phonology. Denis Arnold, Petra Wagner, Bernd Möbius |
INTERSPEECH | 2 |
| 2012 | Gaze Patterns in Turn-TakingabstractOertel C, Wlodarczak M, Edlund J, Wagner P, Gustafson J. Gaze patterns in turn-taking. In: 13th Annual Conference of the International Speech Communication Association 2012 (INTERSPEECH 2012). Red Hook, NY: Curran; 2013: 2243-2246. Catharine Oertel, Marcin Wlodarczak, Jens Edlund, Petra Wagner, Joakim Gustafson |
INTERSPEECH | 4 |
| 2012 | Disentangling lexical, morphological, syntactic and semantic influences on German prominence - Evidence from a production studyabstractSamlowski B, Wagner P, Möbius B. Disentangling lexical, morphological, syntactic and semantic influences on German prominence – Evidence from a production study. In: Proceedings of Interspeech 2012. 2012: 2406-2409. Barbara Samlowski, Petra Wagner, Bernd Möbius |
INTERSPEECH | 2 |
| 2012 | Objective, Subjective and Linguistic Roads to Perceptual Prominence - How are they compared and why?abstractProsodic prominence denotes the perceptual salience of linguistic units. There exists no agreement on (1) ade-quate methods for its subjective measurement, (2) its ob-jective acoustic correlates and (3) its relationship to lin-guistic structure. A traditional approach for evaluating any of these descriptive layers is an inter-level compari-son, e.g. between a perceptual and an acoustic model of prominence. However, (1) there exists no standard pro-cedure for such a comparison, and (2) such a comparison is misleading if both layers are expected to be symmetri-cal, given the neglected influence of linguistic top-down expectancies. We propose an evaluation procedure for prominence models relying on tripartite correlations of perception, its acoustic correlates and linguistic expecta-tions. We suggest a novel correlation metric and test its usefulness on a prosodic corpus of German. Index Terms: prominence, evaluation, prosody 1. Introduction: Model Petra Wagner, Fabio Tamburini, Andreas Windmann |
INTERSPEECH | 1 |
| 2012 | Temporal entrainment in overlapped speech: Cross-linguistic studyabstractWlodarczak M, Simko J, Wagner P. Temporal entrainment in overlapped speech: Cross-linguistic study. In: 13th Annual Conference of the International Speech Communication Association 2012 (INTERSPEECH 2012). Vol. 1. Red Hook, NY: Curran; 2013: 614-617. Marcin Wlodarczak, Juraj Simko, Petra Wagner |
INTERSPEECH | 3 |
| 2011 | Dynamic perception-production oscillation model in human-machine communicationabstractThe goal of the present article is to introduce a new concept of a perception-production timing model in human-machine communication. The model implements a low-level cognitive timing and coordination mechanism. The basic element of the model is a dynamic oscillator capable of tracking reoccurring events in time. The organization of the oscillators in a network is being referred to as the Dynamic Perception-Production Oscillation Model (DPPOM). The DPPOM is largely based on findings in psychological and phonetic experiments on timing in speech perception and production. It consists of two sub-systems, a perception sub-system and a production sub-system. The perception sub-system accounts for information clustering in an input sequence of events. The production sub-system accounts for speech production rhythmically entrained to the input sequence. We propose a system architecture integrating both sub-systems, providing a flexible mechanism for perception-production timing in dialogues. The model's functionality was evaluated in two experiments. Igor Jauk, Ipke Wachsmuth, Petra Wagner |
ICMI | 3 |
| 2011 | Comparing Word and Syllable Prominence Rated by Naïve ListenersabstractArnold D, Möbius B, Wagner P. Comparing word and syllable prominence rated by naive listeners. In: Proceedings of Interspeech 2011. 2011: 1877-1880. Denis Arnold, Bernd Möbius, Petra Wagner |
INTERSPEECH | 3 |
| 2011 | 'Are You Sure You're Paying Attention?' - 'Uh-Huh' Communicating Understanding as a Marker of AttentivenessabstractBuschmeier H, Malisz Z, Wlodarczak M, Kopp S, Wagner P. 'Are you sure you're paying attention?' – 'Uh-huh'. Communicating understanding as a marker of attentiveness. In: Proceedings of INTERSPEECH 2011. International Speech Communication Association; 2011: 2057-2060. Hendrik Buschmeier, Zofia Malisz, Marcin Wlodarczak, Stefan Kopp, Petra Wagner |
INTERSPEECH | 5 |
| 2011 | Comparing Syllable Frequencies in Corpora of Written and Spoken LanguageabstractIn the study, various German language corpora were compared in order to discover the extent to which syllable frequencies remain stable across different contexts and modalities. Although considerable differences in relative frequency were found among the more common syllables, rank numbers proved to be more robust. Variation across corpora was mostly due to vocabulary characteristics of particular corpus domains rather than to systematic differences between spoken and written language. The results indicate that syllable frequencies in written corpora can be taken as a rough estimate for their frequency in spoken language. Index Terms: syllabary, syllable frequencies, spoken and written language corpora Barbara Samlowski, Bernd Möbius, Petra Wagner |
INTERSPEECH | 3 |
| 2011 | Using Prominence Detection to Generate Acoustic Feedback in Tutoring ScenariosabstractRobots interacting with humans need to understand actions and make use of language in social interactions. Research on infant development has shown that language helps the learner to structure visual observations of action. This acoustic information typically in the form of narration overlaps with action sequences and provides infants with a bottom-up guide to find structure within them. This concept has been introduced as acoustic packaging by Hirsh-Pasek and Golinkoff. We developed and integrated a prominence detection module in our acoustic packaging system to detect semantically relevant information linguistically\nhighlighted by the tutor. Evaluation results on speech data from adult-infant interactions show a significant agreement with human raters. Furthermore a first approach based on acoustic packages which uses the prominence detection results to generate acoustic feedback is presented.\n\nIndex Terms: prominence, multimodal action segmentation,\nhuman robot interaction, feedback Lars Schillingmann, Petra Wagner, Christian Munier, Britta Wrede, Katharina J. Rohlfing |
INTERSPEECH | 2 |
| 2011 | Prominence-Based Prosody Prediction for Unit Selection Speech SynthesisabstractThis paper describes the development and evaluation of a\nprosody prediction module for unit selection speech synthesis that is based on the notion of perceptual prominence. We outline the design principles of the module and describe its implementation in the Bonn Open Synthesis System (BOSS).\nMoreover, we report results of perception experiments that\nhave been conducted in order to evaluate prominence\nprediction. The paper is concluded by a general discussion of the approach and a sketch of perspectives for further work. Andreas Windmann, Igor Jauk, Fabio Tamburini, Petra Wagner |
INTERSPEECH | 4 |
| 2009 | Assessing a speaker for fast speech in unit selection speech synthesisabstractMoers D, Wagner P. Assessing a speaker for fast speech in unit selection speech synthesis. In: Proceedings of Interspeech. 2009: 2071-2074. Donata Moers, Petra Wagner |
INTERSPEECH | 2 |
| 2009 | Paper 8003 was not available at the time of publication oral presentation of poster papers no time to lose? time shrinking effects enhance the impression of rhythmic "isochrony" and fast speech rate
Petra Wagner, Andreas Windmann |
INTERSPEECH | 1 |
| 2007 | On automatic prominence detection for GermanabstractPerceptual prominence is an important indicator of a word's and syllable's lexical, syntactic, \nsemantic and pragmatic status in a discourse. Its automatic annotation would be a valuable \nenrichment of large databases used in unit selection speech synthesis and speech recognition. \nWhile much research has been carried out on the interaction between prominence and \nacoustic factors, little progress has been made in its automatic annotation. Previous \napproaches to German relied on linguistic features in prominence detection, but a purely \nacoustic method would be advantageous. We applied an algorithm to German data that had \nbeen previously used for English and Italian. Both the algorithm and the data annotation \nencode prominence as a continuous rather than a categorical parameter. First results are \nencouraging, but again show that prominence perception relies on linguistic expectancies as \nwell as acoustic patterns. Also, our results further strengthen the view that force accents are a \nmore reliable cue to prominence than pitch accents in German. Fabio Tamburini, Petra Wagner |
INTERSPEECH | 2 |
| 2005 | Great expectations - introspective vs. perceptual prominence ratings and their acoustic correlatesabstractIn order to gain knowledge about the interaction between top-down expectations of listeners concerning prosodic prominence and its acoustic correlates, two exploratory empirical studies were carried out. First, native and non-native subjects rated prominences of speech read at normal and very fast - prosodically very different - speech. Later, these ratings were compared with introspective prominence ratings of different listeners. First results indicate a major influence of the introspection on prominence ratings, especially if acoustic cues are difficult to interpret, as it is the case in very fast speech. Compared to native subjects, non-natives rely less on their introspection and more on the acoustics. Petra Wagner |
INTERSPEECH | 1 |
| 2004 | Bonntempo-corpus and bonntempo-tools: a database for the study of speech rhythm and rateabstractWork is currently being carried out on a speech database
constructed in order to study speech rhythm in connection with speech rate. The database, BonnTempo-Corpus, and the Praat based analysis tools, BonnTempo-Tools, are a powerful
instrument for examining various aspects of recently proposed rhythm measures (e.g. %V, C, nPVI, rPVI, etc.) in relation to speech rate among a wide range of languages and speakers.
First observations pose new problems on traditionally not well classifiable languages like Czech. Volker Dellwo, Bianca Aschenberner, Petra Wagner, Jana Dancovicova, Ingmar Steiner |
INTERSPEECH | 3 |
| 2001 | Speech synthesis development made easy: the bonn open synthesis systemabstractThis paper describes a new open source architecture for unit-selection based speech synthesis called BOSS (Bonn Open Synthesis System). It is built up modularly, with communications between modules taking place in a fixed format. This makes the addition, deletion and substitution of modules very easy. The strict separation between data and algorithms allows for the simple creation of new speech corpora for different domains and languages. 1. Esther Klabbers, Karlheinz Stöber, Raymond N. J. Veldhuis, Petra Wagner, Stefan Breuer |
INTERSPEECH | 4 |
| 1999 | Synthesis by word concatenationabstractStöber K, Portele T, Wagner P, Hess W. Synthesis by Word Concatenation. In: Proceedings of Interspeech 1999. Vol 2. Budapest, Hungary; 1999: 619-622. Karlheinz Stöber, Thomas Portele, Petra Wagner, Wolfgang Hess |
EUROSPEECH | 3 |