Katharina Zahner-Ritter

dblp:173/6644 · also Katharina Zahner · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-1954-6436ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Are You Being Sarcastic? Prosodic Cues to Irony Perception in German
Sophia Fünfgeld, Angelika Braun, Katharina Zahner-Ritter
INTERSPEECH3
2024 The use of Active Learning systems for stimulus selection and response modelling in perception experiments
abstract
To study the role of perceptual cues on categorization and decision making, participants are typically tested in (perception) experiments with a fixed set of randomized or pseudo-randomized trials. In linguistics and psycholinguistics, for instance, studies often investigate the relative weighting of different cues for a linguistic contrast (e.g., intonation vs. word order). For categorization beyond the segmental level (e.g., /p/ vs. /b/), it is important to establish that results generalise to different words or sentences, which necessitates the use of a range of different items. This may limit the number of conditions (cues and cue combinations) that can be sensibly tested in the same experiment. We show that Active Learning (AL) systems provide a solution: Since stimulus selection is informed by the system's learning mechanism (presenting obvious conditions less often than uncertain conditions), they allow for efficient testing of numerous conditions and different items in the same experiment. In this paper, we compared two weighting approaches (probability-based vs. regression-based) to model the outcome of three simulated scenarios with three binary factors each. Results show that valid results (i.e., little error between predicted values and the actual responses at the end of the experiment) are obtained after about half of the trials of an original psycholinguistic experiment we replicated. For simulations with interactions between factors, the regression-based approach performed better. Our findings bear implications for the application of AL in psycholinguistic research (extraction of cue weights, inferential statistics, and a stopping criterion during an on-going experiment), which we will discuss.
Marieke Einfeldt, Rita Sevastjanova, Katharina Zahner-Ritter, Ekaterina Kazak, Bettina Braun
Comput. Speech Lang.3
2023 Speech Enhancement Patterns in Human-Robot Interaction: A Cross-Linguistic Perspective
abstract
This paper presents the results of the human-robot interaction (HRI) study with German native speakers addressing the robot in their L1 and in L2 English. The aim of the experiment is to test the strategies of providing clarifications when talking to the voice assistant in a task involving teaching complex vocabulary. The analyses is based on spectral (F1, F2, and mean F0) and temporal (vowel length) features excerpted from the target words. With reference to a theoretical framework of hyperarticulation and hypoarticulation, these acoustic measures were compared across the iterations of the target words (first vs. second iteration). Results showed that participants, when asked for clarification by an inanimate interlocutor, do not hyperarticulate, but try to preserve the surface representation of target words across the iterations. These findings suggest that acoustic characteristics of clarifications directed to voice assistants differ from the ones directed to human interlocutors.
Jacek Kudera, Katharina Zahner-Ritter, Jakob Engel, Nathalie Elsässer, Philipp Hutmacher, Carolin Worstbrock
INTERSPEECH2
2021 Testing Acoustic Voice Quality Classification Across Languages and Speech Styles
abstract
Many studies relate acoustic voice quality measures to perceptual classification. We extend this line of research by training a classifier on a balanced set of perceptually annotated voice quality categories with high inter-rater agreement, and test it on speech samples from a different language and on a different speech style. Annotations were done on continuous speech from different laboratory settings. In Experiment 1, we trained a random forest with Standard Chinese and German recordings labelled as modal, breathy, or glottalized. The model had an accuracy of 78.7% on unseen data from the same sample (most important variables were harmonics-to-noise ratio, cepstral-peak prominence, and H1-A2). This model was then used to classify data from a different language (Icelandic, Experiment 2) and to classify a different speech style (German infant-directed speech (IDS), Experiment 3). Cross-linguistic generalizability was high for Icelandic (78.6% accuracy), but lower for German IDS (71.7% accuracy). Accuracy of recordings of adult-directed speech from the same speakers as in Experiment 3 (77%, Experiment 4) suggests that it is the special speech style of IDS, rather than the recording setting that led to lower performance. Results are discussed in terms of efficiency of coding and generalizability across languages and speech styles.
Bettina Braun, Nicole Dehé, Marieke Einfeldt, Daniela Wochner, Katharina Zahner-Ritter
Interspeech5
2021 Reliable Estimates of Interpretable Cue Effects with Active Learning in Psycholinguistic Research
Marieke Einfeldt, Rita Sevastjanova, Katharina Zahner-Ritter, Ekaterina Kazak, Bettina Braun
Interspeech3
2021 In-Group Advantage in the Perception of Emotions: Evidence from Three Varieties of German
abstract
Various studies on the perception of vocally expressed emo-tions have shown that recognition rates are higher if speaker and listener belong to the same cultural or linguistic group. This so-called in-group advantage is commonly attributed to prosodic differences in the expression of emotion across groups. Evidence comes mostly from using cross-linguistic and/or cross-cultural study designs. Previous research suggests that varieties of German differ in their use of prosody and can be discriminated based on prosodic features alone. In this paper, we tested whether emotion recognition rates differ across varieties of German: Listeners from three dialectal areas (Hamburg, Vienna, Zurich) identified emotions on semantically neutral sen-tences (choosing between anger, happiness, relief, surprise or “other”), spoken by actors from the three regions. Correctness rates show that emotions are recognized better if speakers and listeners are native speakers of the same variety. However, further analyses suggest that the in-group advantage does not surface consistently across individual emotions. To explain these results, the prosodic realization of the sentences was tested for interactions between emotion and variety. Here, intensity seemed to differ most across varieties and emotions. Im-portantly, we show that the in-group advantage extends from cultural groups to dialectal groups of a language.
Moritz Jakob, Bettina Braun, Katharina Zahner-Ritter
Interspeech3
2018 Truncation and Compression in Southern German and Australian English
Jenny Yu, Katharina Zahner-Ritter
INTERSPEECH2
2018 The Distribution and Prosodic Realization of Verb Forms in German Infant-Directed Speech
Bettina Braun, Katharina Zahner-Ritter
LREC2
2017 Similar Prosodic Structure Perceived Differently in German and English
abstract
English and German have similar prosody, but their speakers realize some pitch falls (not rises) in subtly different ways. We here test for asymmetry in perception. An ABX discrimination task requiring F0 slope or duration judgements on isolated vowels revealed no cross-language difference in duration or F0 fall discrimination, but discrimination of rises (realized similarly in each language) was less accurate for English than for German listeners. This unexpected finding may reflect greater sensitivity to rising patterns by German listeners, or reduced sensitivity by English listeners as a result of extensive exposure to phrase-final rises ("uptalk") in their language.
Heather Kember, Ann-Kathrin Grohe, Katharina Zahner-Ritter, Bettina Braun, Andrea Weber, Anne Cutler
INTERSPEECH3
2017 Mind the Peak: When Museum is Temporarily Understood as Musical in Australian English
abstract
Intonation languages signal pragmatic functions (e.g. information structure) by means of different pitch accent types. Acoustically, pitch accent types differ in the alignment of pitch peaks (and valleys) in regard to stressed syllables, which makes the position of pitch peaks an unreliable cue to lexical stress (even though pitch peaks and lexical stress often coincide in intonation languages). We here investigate the effect of pitch accent type on lexical activation in English. Results of a visual-world eye-tracking study show that Australian English listeners temporarily activate SWW-words ( musical) if presented with WSW-words ( museum) with early-peak accents (H+!H*), compared to medial-peak accents (L+H*). Thus, in addition to signalling pragmatic functions, the alignment of tonal targets immediately affects lexical activation in English.
Katharina Zahner-Ritter, Heather Kember, Bettina Braun
INTERSPEECH1
2015 Pitch accent distribution in German infant-directed speech
abstract
Infant-directed speech exhibits slower speech rate, higher pitch and larger f0-excursions than adult-directed speech. Apart from these phonetic properties established in many languages, little is known on the intonational phonological structure in individual languages, i.e. pitch accents and boundary tones and their frequency distribution. Here, we investigated the intonation of infant-directed speech in German. We extracted all turns from the CHILDES database directed towards infants younger than one year (n=585). Two annotators labeled pitch accents and boundary tones according to the autosegmental-metrical intonation system GToBI. Additionally, the tonal movement surrounding the accentual syllable was analyzed. Main results showed a) that 45 % of the words carried a pitch accent, b) that phrases ending in a low tone were most frequent, c) that H * accents were generally more frequent than L * accents, d) that H*, L+H * and L * are the most frequent pitch accent types in IDS, and e) that a pattern consisting of an accentual low-pitched syllable preceded by a low tone and followed by a rise or a high tone constitutes the most frequent single pattern. The analyses reveal that the IDS intonational properties lead to a speech style with many tonal alternations, particularly in the vicinity of accented syllables. Index Terms: intonation, infant-directed speech, pitch accent
Katharina Zahner-Ritter, Muna Pohl, Bettina Braun
INTERSPEECH1