EDBT 2026 Demo / reviewers in the wild / expert
Bernd Möbius
dblp:79/4327
· DBLP profile ↗
72ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0003-3065-9984ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 66 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 63 · 5 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automatic Prediction of Child Speech Fluency with Game-Based Data from German Preschoolers
Valentin Kany, Bernd Möbius, Jürgen Trouvain |
LREC | 2 |
| 2025 | It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text SystemsabstractIuliia Zaitova, Badr M. Abdullah, Wei Xue, Dietrich Klakow, Bernd Möbius, Tania Avgustinova. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Iuliia Zaitova, Badr Abdullah, Dietrich Klakow, Bernd Möbius, Tania Avgustinova |
ACL (1) | 5 |
| 2025 | Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification
Badr Abdullah, Matthew Baas, Bernd Möbius, Dietrich Klakow |
INTERSPEECH | 3 |
| 2025 | The Effect of Word Predictability on Spoken Cross-Language Intelligibility
Iuliia Zaitova, Bernd Möbius |
INTERSPEECH | 3 |
| 2024 | Towards a better understanding of receptive multilingualism: listening conditions and priming effects
Ivan Yuen, Bernd Möbius |
INTERSPEECH | 3 |
| 2024 | Cross-Linguistic Intelligibility of Non-Compositional Expressions in Spoken Context
Iuliia Zaitova, Irina Stenger, Tania Avgustinova, Bernd Möbius, Dietrich Klakow |
INTERSPEECH | 5 |
| 2024 | Forms, factors and functions of phonetic convergence: Editorial
Elisa Pellegrino, Volker Dellwo, Jennifer S. Pardo, Bernd Möbius |
Speech Commun. | 4 |
| 2023 | An Information-Theoretic Analysis of Self-supervised Discrete Representations of Speech
Badr Abdullah, Mohammed Maqsood Shaik, Bernd Möbius, Dietrich Klakow |
INTERSPEECH | 3 |
| 2023 | Cross-linguistic Emotion Perception in Human and TTS Voices
Iona Gessinger, Michelle Cohn, Benjamin R. Cowan, Georgia Zellou, Bernd Möbius |
INTERSPEECH | 5 |
| 2022 | Integrating Form and Meaning: A Multi-Task Learning Model for Acoustic Word Embeddings
Badr Abdullah, Bernd Möbius, Dietrich Klakow |
INTERSPEECH | 2 |
| 2022 | Cross-Cultural Comparison of Gradient Emotion Perception: Human vs. Alexa TTS Voices
Iona Gessinger, Michelle Cohn, Georgia Zellou, Bernd Möbius |
INTERSPEECH | 4 |
| 2022 | Modeling the Impact of Syntactic Distance and Surprisal on Cross-Slavic Text ComprehensionabstractWe focus on the syntactic variation and measure syntactic distances between nine Slavic languages (Belarusian, Bulgarian, Croatian, Czech, Polish, Slovak, Slovene, Russian, and Ukrainian) using symmetric measures of insertion, deletion and movement of syntactic units in the parallel sentences of the fable “The North Wind and the Sun”. Additionally, we investigate phonetic and orthographic asymmetries between selected languages by means of the information theoretical notion of surprisal. Syntactic distance and surprisal are, thus, considered as potential predictors of mutual intelligibility between related languages. In spoken and written cloze test experiments for Slavic native speakers, the presented predictors will be validated as to whether variations in syntax lead to a slower or impeded intercomprehension of Slavic texts. Irina Stenger, Philip Georgis, Tania Avgustinova, Bernd Möbius, Dietrich Klakow |
LREC | 4 |
| 2021 | Do Acoustic Word Embeddings Capture Phonological Similarity? An Empirical StudyabstractSeveral variants of deep neural networks have been successfully employed for building parametric models that project variable-duration spoken word segments onto fixed-size vector representations, or acoustic word embeddings (AWEs). However, it remains unclear to what degree we can rely on the distance in the emerging AWE space as an estimate of word-form similarity. In this paper, we ask: does the distance in the acoustic embedding space correlate with phonological dissimilarity? To answer this question, we empirically investigate the performance of supervised approaches for AWEs with different neural architectures and learning objectives. We train AWE models in controlled settings for two languages (German and Czech) and evaluate the embeddings on two tasks: word discrimination and phonological similarity. Our experiments show that (1) the distance in the embedding space in the best cases only moderately correlates with phonological distance, and (2) improving the performance on the word discrimination task does not necessarily yield models that better reflect word phonological similarity. Our findings highlight the necessity to rethink the current intrinsic evaluations for AWEs. Badr Abdullah, Marius Mosbach, Iuliia Zaitova, Bernd Möbius, Dietrich Klakow |
Interspeech | 4 |
| 2021 | Take a Breath: Respiratory Sounds Improve Recollection in Synthetic Speech
Mikey Elmers, Raphael Werner, Beeke Muhlack, Bernd Möbius, Jürgen Trouvain |
Interspeech | 4 |
| 2021 | Phonetic Distance and Surprisal in Multilingual Priming: Evidence from Slavic
Jacek Kudera, Philip Georgis, Bernd Möbius, Tania Avgustinova, Dietrich Klakow |
Interspeech | 3 |
| 2021 | Revisiting Recall Effects of Filler Particles in German and English
Beeke Muhlack, Mikey Elmers, Heiner Drenhaus, Jürgen Trouvain, Marjolein van Os, Raphael Werner, Margarita Ryzhova, Bernd Möbius |
Interspeech | 8 |
| 2021 | Inhalations in Speech: Acoustic and Physiological Characteristics
Raphael Werner, Susanne Fuchs, Jürgen Trouvain, Bernd Möbius |
Interspeech | 4 |
| 2021 | Phonetic accommodation to natural and synthetic voices: Behavior of groups and individuals in speech shadowingabstractThe present study investigates whether native speakers of German phonetically accommodate to natural and synthetic voices in a shadowing experiment. We aim to determine whether this phenomenon, which is frequently found in HHI, also occurs in HCI involving synthetic speech. The examined features pertain to different phonetic domains: allophonic variation, schwa epenthesis, realization of pitch accents, word-based temporal structure and distribution of spectral energy. On the individual level, we found that the participants converged to varying subsets of the examined features, while they maintained their baseline behavior in other cases or, in rare instances, even diverged from the model voices. This shows that accommodation with respect to one particular feature may not predict the behavior with respect to another feature. On the group level, the participants of the natural condition converged to all features under examination, however very subtly so for schwa epenthesis. The synthetic voices, while partly reducing the strength of effects found for the natural voices, triggered accommodating behavior as well. The predominant pattern for all voice types was convergence during the interaction followed by divergence after the interaction. Iona Gessinger, Eran Raveh, Ingmar Steiner, Bernd Möbius |
Speech Commun. | 4 |
| 2020 | Cross-Domain Adaptation of Spoken Language Identification for Related Languages: The Curious Case of Slavic LanguagesabstractState-of-the-art spoken language identification (LID) systems, which are based on end-to-end deep neural networks, have shown remarkable success not only in discriminating between distant languages but also between closely-related languages or even different spoken varieties of the same language. However, it is still unclear to what extent neural LID models generalize to speech samples with different acoustic conditions due to domain shift. In this paper, we present a set of experiments to investigate the impact of domain mismatch on the performance of neural LID systems for a subset of six Slavic languages across two domains (read speech and radio broadcast) and examine two low-level signal descriptors (spectral and cepstral features) for this task. Our experiments show that (1) out-of-domain speech samples severely hinder the performance of neural LID models, and (2) while both spectral and cepstral features show comparable performance within-domain, spectral features show more robustness under domain mismatch. Moreover, we apply unsupervised domain adaptation to minimize the discrepancy between the two domains in our study. We achieve relative accuracy improvements that range from 9% to 77% depending on the diversity of acoustic conditions in the source domain. Badr Abdullah, Tania Avgustinova, Bernd Möbius, Dietrich Klakow |
INTERSPEECH | 3 |
| 2020 | Differences in Gradient Emotion Perception: Human vs. Alexa Voices
Michelle Cohn, Eran Raveh, Kristin Predeck, Iona Gessinger, Bernd Möbius, Georgia Zellou |
INTERSPEECH | 5 |
| 2020 | Phonetic Accommodation of L2 German Speakers to the Virtual Language Learning Tutor Mirabella
Iona Gessinger, Bernd Möbius, Bistra Andreeva, Eran Raveh, Ingmar Steiner |
INTERSPEECH | 2 |
| 2019 | Phonetic Accommodation in a Wizard-of-Oz Experiment: Intonation and Segments
Iona Gessinger, Bernd Möbius, Bistra Andreeva, Eran Raveh, Ingmar Steiner |
INTERSPEECH | 2 |
| 2019 | Three's a Crowd? Effects of a Second Human on Vocal Accommodation with a Voice Assistant
Eran Raveh, Ingo Siegert, Ingmar Steiner, Iona Gessinger, Bernd Möbius |
INTERSPEECH | 5 |
| 2018 | A change at the helm for Speech Communication
Bernd Möbius |
Speech Commun. | 1 |
| 2017 | Mel-Cepstral Distortion of German Vowels in Different Information Density Contexts
Erika Brandt, Frank Zimmerer, Bistra Andreeva, Bernd Möbius |
INTERSPEECH | 4 |
| 2017 | Shadowing Synthesized Speech - Segmental Analysis of Phonetic Convergence
Iona Gessinger, Eran Raveh, Sébastien Le Maguer, Bernd Möbius, Ingmar Steiner |
INTERSPEECH | 4 |
| 2017 | A Computational Model for Phonetically Responsive Spoken Dialogue Systems
Eran Raveh, Ingmar Steiner, Bernd Möbius |
INTERSPEECH | 3 |
| 2016 | The Perceptual Effect of L1 Prosody Transplantation on L2 Speech: The Case of French Accented GermanabstractResearch has shown that language learners are not only challenged by segmental differences between their native language (L1) and the second language (L2). They also have problems with the correct production of suprasegmental structures, like phone/syllable duration and the realization of pitch. These difficulties often lead to a perceptible foreign accent. This study investigates the influence of prosody transplantation on foreign accent ratings. Syllable duration and pitch contour were transferred from utterances of a male and female German native speaker to utterances of ten French native speakers speaking German. Acoustic measurements show that French learners spoke with a significantly lower speaking rate. As expected, results of a perception experiment judging the accentedness of 1) German native utterances, 2) unmanipulated and 3) manipulated utterances of French learners of German suggest that the transplantation of the prosodic features syllable duration and pitch leads to a decrease in accentedness rating. These findings confirm results found in similar studies investigating prosody transplantation with different L1 and L2 and provide a beneficial technique for (computer-assisted) pronunciation training. Jeanin Jügler, Frank Zimmerer, Jürgen Trouvain, Bernd Möbius |
INTERSPEECH | 4 |
| 2016 | The IFCASL Corpus of French and German Non-native and Native Read Speech
Jürgen Trouvain, Anne Bonneau, Vincent Colotte, Camille Fauth, Dominique Fohr, Denis Jouvet, Jeanin Jügler, Yves Laprie, Odile Mella, Bernd Möbius, Frank Zimmerer |
LREC | 10 |
| 2015 | Linguistic measures of pitch range in slavic and Germanic languagesabstractBased on specific linguistic landmarks in the speech signal, this study investigates pitch level and pitch span differences in English, German, Bulgarian and Polish. The analysis is based on 22 speakers per language (11 males and 11 females). Linear mixed models were computed that include various linguistic measures of pitch level and span, revealing characteristic differences across languages and between language groups. Pitch level appeared to have significantly higher values for the female speakers in the Slavic than the Germanic group. The male speakers showed slightly different results, with only the Polish speakers displaying significantly higher mean values for pitch level than the German males. Overall, the results show that the Slavic speakers tend to have a wider pitch span than the German speakers. But for the linguistic measure, namely for span between the initial peaks and the non-prominent valleys, we only find the difference between Polish and German speakers. We found a flatter intonation contour in German than in Polish, Bulgarian and English male and female speakers and differences in the frequency of the landmarks between languages. Concerning “speaker liveliness” we found that the speakers from the Slavic group are significantly livelier than the speakers from the Germanic group. Bistra Andreeva, Bernd Möbius, Grazyna Demenko, Frank Zimmerer, Jeanin Jügler |
INTERSPEECH | 2 |
| 2015 | The effect of high-variability training on the perception and production of French stops by German native speakersabstractWe investigated the effect of high-variability training (HVT) on the production and perception of French bilabial voiced and voiceless stops by German native speakers. Stop consonants in the two languages differ with respect to several articulatory and acoustic features. German learners of French (Experiment Group) trained the perception of word-initial bilabial stops spoken by six French native speakers using identification tests, whereas subjects of a Control Group did not perform a training. Additional perception and production tests of French words including bilabial, alveolar, and velar stops in all word positions were performed to capture the impact of HVT. Subjects were found to be quite good at distinguishing voiced and voiceless stops. However, voiceless stops received lower correctness scores than voiced ones and subjects of the Experiment group were able to further increase their scores after training. Results for production are mirror-inverted showing that subjects of the Experiment Group successfully produced longer negative VOT values but did not show an improvement for voiceless stops. Jeanin Jügler, Frank Zimmerer, Bernd Möbius, Christoph Draxler |
INTERSPEECH | 3 |
| 2015 | Exploring the relationship between intonation and the lexicon: Evidence for lexicalised storage of intonation
Katrin Schweitzer, Michael Walsh 0001, Sasha Calhoun, Hinrich Schütze, Bernd Möbius, Antje Schweitzer, Grzegorz Dogil |
Speech Commun. | 5 |
| 2014 | Differences of pitch profiles in Germanic and slavic languagesabstractThis study investigates cross-language differences in pitch range and variation in four languages from two language groups: English and German (Germanic) and Bulgarian and Polish (Slavic). The analysis is based on large multi-speaker corpora (48 speakers for Polish, 60 for each of the other three languages). Linear mixed models were computed that include various distributional measures of pitch level, span and variation, revealing characteristic differences across languages and between language groups. A classification experiment based on the relevant parameter measures (span, kurtosis and skewness values for pitch distributions for each speaker) succeeded in separating the language groups. Index Terms: pitch range, pitch variation, cross-language Bistra Andreeva, Grazyna Demenko, Bernd Möbius, Frank Zimmerer, Jeanin Jügler, Magdalena Oleskowicz-Popiel |
INTERSPEECH | 3 |
| 2014 | Designing a Bilingual Speech Corpus for French and German Language Learners: a Two-Step Process
Camille Fauth, Anne Bonneau, Frank Zimmerer, Jürgen Trouvain, Bistra Andreeva, Vincent Colotte, Dominique Fohr, Denis Jouvet, Jeanin Jügler, Yves Laprie, Odile Mella, Bernd Möbius |
LREC | 12 |
| 2013 | Effects of lexical class and lemma frequency on German homographsabstractGerman demonstrative pronouns, relative pronouns, and definite articles are segmentally identical but differ strongly in the frequency with which they appear. We examined the production of five such particles in a reading task. In a comparison of orthographically identical word pairs belonging to different lexical classes we found small but significant differences in word and vowel duration, prominence, and spectral similarity. Three of the particles in particular tended to be longer and more prominent when they occurred as demonstrative or relative articles than when they were assigned their usual role as definite articles. Index Terms: lemma frequency, duration, prominence 1. Barbara Samlowski, Petra Wagner, Bernd Möbius |
INTERSPEECH | 3 |
| 2012 | Obtaining prominence judgments from naïve listeners - Influence of rating scales, linguistic levels and normalisationabstractA frequently replicated finding is that higher frequency words tend to be shorter and contain more strongly reduced vowels. However, little is known about potential differences in the articulatory gestures for high vs. low frequency words. The present study made use of electromagnetic articulography to investigate the production of two German vowels, [i] and [a], embedded in high and low frequency words. We found that word frequency differently affected the production of [i] and [a] at the temporal as well as the gestural level. Higher frequency of use predicted greater acoustic durations for long vowels; reduced durations for short vowels; articulatory trajectories with greater tongue height for [i] and more pronounced downward articulatory trajectories for [a]. These results show that the phonological contrast between short and long vowels is learned better with experience, and challenge both the Smooth Signal Redundancy Hypothesis and current theories of German phonology. Denis Arnold, Petra Wagner, Bernd Möbius |
INTERSPEECH | 3 |
| 2012 | Describing the development of intonational categories using a target-oriented parametric approach
Britta Lintfert, Bernd Möbius |
INTERSPEECH | 2 |
| 2012 | Disentangling lexical, morphological, syntactic and semantic influences on German prominence - Evidence from a production studyabstractSamlowski B, Wagner P, Möbius B. Disentangling lexical, morphological, syntactic and semantic influences on German prominence – Evidence from a production study. In: Proceedings of Interspeech 2012. 2012: 2406-2409. Barbara Samlowski, Petra Wagner, Bernd Möbius |
INTERSPEECH | 3 |
| 2011 | Comparing Word and Syllable Prominence Rated by Naïve ListenersabstractArnold D, Möbius B, Wagner P. Comparing word and syllable prominence rated by naive listeners. In: Proceedings of Interspeech 2011. 2011: 1877-1880. Denis Arnold, Bernd Möbius, Petra Wagner |
INTERSPEECH | 2 |
| 2011 | A Parametric Approach to Intonation Acquisition Research: Validation on Child-Directed Speech DataabstractThis paper validates a parametric approach to intonation acquisition research [1] using child-directed speech data. An advantage of this approach is that it can be used for studying child speech as well as adult speech. Within the field of prosody acquisition it reconciles independent approaches to child prosody with ToBI-based approaches. In this paper we substantiate this claim by showing that clusters of parameterized contours obtained from German child-directed speech correlate with GToBI(S) categories, and by elaborating how, alternatively, the parameters can be mapped to properties that are relevant in independent approaches. Index Terms: intonation, F0 parameterization, clustering, child-directed speech Britta Lintfert, Antje Schweitzer, Bernd Möbius |
INTERSPEECH | 3 |
| 2011 | Comparing Syllable Frequencies in Corpora of Written and Spoken LanguageabstractIn the study, various German language corpora were compared in order to discover the extent to which syllable frequencies remain stable across different contexts and modalities. Although considerable differences in relative frequency were found among the more common syllables, rank numbers proved to be more robust. Variation across corpora was mostly due to vocabulary characteristics of particular corpus domains rather than to systematic differences between spoken and written language. The results indicate that syllable frequencies in written corpora can be taken as a rough estimate for their frequency in spoken language. Index Terms: syllabary, syllable frequencies, spoken and written language corpora Barbara Samlowski, Bernd Möbius, Petra Wagner |
INTERSPEECH | 2 |
| 2011 | Reaction Time and Decision Difficulty in the Perception of IntonationabstractAn experiment was carried out to test the Categorical Percep-tion as well as possible Perceptual Magnet Effects in the two boundary tone categories L % and H % in German, correspond-ing to statement vs. question interpretation, respectively. Addi-tionally, reaction times (RT) were logged during all subtests to see if they support the results. Analyses revealed that RTs al-ways increased with rising difficulty of the perceptual task, and decreased when the decision process was easy. Task-specific re-sults showed that RT also correlated with the number of possible answers during a perceptual decision, i.e. more answer alterna-tives resulted in longer RT. Furthermore, female subjects gener-ally reacted faster during all perceptual tasks, although this did not necessarily correlate with the accuracy of the results. Nev-ertheless, the results confirmed the usefulness of RT to support the analyses and the interpretation of perceptual data. Index Terms: prosody perception, intonation, categorical per-ception, perceptual magnet effect, reaction time Katrin Schneider, Grzegorz Dogil, Bernd Möbius |
INTERSPEECH | 3 |
| 2010 | Frequency of occurrence effects on pitch accent realisationabstractThis paper presents the results of a corpus study which examines the impact of frequency of occurrence of accented words on the realisation of pitch accents. In particular, statistical analyses explore this influence on pitch accent range and alignment. The results indicate a significant effect of frequency of occurrence on the relative height of L*H and H*L pitch accents and an also significant but more subtle effect on the alignment of L*H accents. Index Terms: exemplar theory, pitch accents, frequency effects Katrin Schweitzer, Michael Walsh 0001, Bernd Möbius, Hinrich Schütze |
INTERSPEECH | 3 |
| 2010 | Specification in context - devoicing processes in Polish, French, american English and German sonorantsabstractThis study investigates voicing properties of Polish, French, American English and German sonorant consonants, particularly rhotics. The analysis was conducted on four speech databases recorded by professional speakers. The term voicing profiles used in this article refers to the frame by frame voicing status of the sonorants, which was obtained by automatic measurements of fundamental frequency values and extraction of consonantal features. Results show resyllabification processes in Polish and French obstruent liquid clusters in word final positions, as well as contextual effects on devoicing in word initial and word medial American English and German obstruent sonorant clusters. Index Terms: voicing, sonorant, rhotic, speech database, Polish, French, American English, German. Jagoda Sieczkowska, Bernd Möbius, Grzegorz Dogil |
INTERSPEECH | 2 |
| 2009 | Frequency Matters: Pitch Accents and Information Status
Katrin Schweitzer, Michael Walsh 0001, Bernd Möbius, Arndt Riester, Antje Schweitzer, Hinrich Schütze |
EACL | 3 |
| 2009 | German boundary tones show categorical perception and a perceptual magnet effect when presented in different contextsabstractThe experiment presented in this paper examines categorical perception as well as the perceptual magnet effect in German boundary tones, taking also context information into account. The test phrase is preceded by different context sentences that are assumed to affect the location of the category boundary in the stimulus continuum between the low and the high boundary tone. Results provide evidence for the existence of a low and a high boundary tone in German, corresponding to statement versus question interpretation, respectively. Furthermore, in contrast to previous findings, a prototype was found not only in the Katrin Schneider, Grzegorz Dogil, Bernd Möbius |
INTERSPEECH | 3 |
| 2009 | Experiments on automatic prosodic labelingabstractThis paper presents results from experiments on automatic prosodic labeling. Using the WEKA machine learning soft-ware [1], classifiers were trained to determine for each syllable in a speech database of a male speaker its pitch accent and its boundary tone. Pitch accents and boundaries are according to the GToBI(S) dialect, with slight modifications. Classification was based on 35 attributes involving PaIntE F0 parametrization [2] and normalized phone durations, but also some phonologi-cal information as well as higher-linguistic information. Several classification algorithms yield results of approx. 78 % accuracy on the word level for pitch accents, and approx. 88 % accuracy on the word level for phrase boundaries, which compare very well to results of other studies. The classifiers generalize to similar data of a female speaker in that they perform equally well as classifiers trained directly on the female data. Index Terms: perception of prosody, prosodic labeling, F0 parametrization Antje Schweitzer, Bernd Möbius |
INTERSPEECH | 2 |
| 2009 | Voicing profile of Polish sonorants: [r] in obstruent clustersabstractThis study aims at defining and analyzing voicing profile of Polish sonorant [r] showing the variability of its realizations depending on segmental and prosodic position. Voicing profile is defined as the frame-by-frame voicing status of a speech sound in continuous speech. Word-final devoicing of sonorants is shortly reviewed and analyzed in terms of the conducted corpus-based investigation. We used automatic tools to extract consonants’ features, F0 values and obtain voicing profile. The results show that liquid [r] devoice word and syllable finally, particularly with left voiceless stop context. Index Terms: sonorant, liquid, voicing, Polish, speech database. Jagoda Sieczkowska, Bernd Möbius, Antje Schweitzer, Michael Walsh 0001, Grzegorz Dogil |
INTERSPEECH | 2 |
| 2008 | Development and evaluation of Polish speech corpus for unit selection speech synthesis systemsabstractThis paper presents the results of a set of experiments assessing the perceived quality of the Polish version of the BOSS unit selection synthesis system. The experiments aimed to evaluate the potential improvement of synthesis quality by three factors pertaining to corpus structure and coverage as well as levels of corpus annotation. The three factors affecting synthesis quality were (i) manual vs. automatic corpus annotation, (ii) coverage of CVC triphones in rich intonational patterns, and (iii) coverage of complex consonant clusters. Results indicate that a manual correction of automatic annotations enhances synthesis quality. Increased coverage of CVC sequences and consonant clusters also improved the perceived synthesis quality, but the effect was smaller than anticipated. Grazyna Demenko, Jolanta Bachan, Bernd Möbius, Katarzyna Klessa, Marcin Szymanski, Stefan Grocholewski |
INTERSPEECH | 3 |
| 2008 | Examining pitch-accent variability from an exemplar-theoretic perspective
Michael Walsh 0001, Katrin Schweitzer, Bernd Möbius, Hinrich Schütze |
INTERSPEECH | 3 |
| 2007 | The influence of vowel quality features on peak alignmentabstractThis study continues an approach that uses a unit selection corpus in order to investigate aspects of the phonetic realization of tonal categories. The focus lies on the peak position of German H*L pitch accents, specifically on the question of whether it is influenced by vowel quality. It is confirmed that vowel backness does not affect peak alignment at all. The distinction between tense and lax vowels initially promises to be relevant, as the H*L peaks seemingly occur significantly earlier in lax vowels. The effect is however demonstrated to be caused by the far greater number of lax vowels in the closed syllables found in the corpus. Finally, the feature of vowel height is revealed to be a significant factor (peaks are aligned latest in high vowels, earliest in low vowels). Various parameters (e.g., syllable structure, position in the phrase) are examined for interactions, but cannot account for the effect. While vowel height correlates with vowel duration, vowel duration itself does not influence peak position. The only possible explanation found involves peak height, which is intrinsically higher in high vowels, thus it may require more time to reach the peak. Index Terms: peak alignment, vowel height, unit selection corpus, German Matthias Jilka, Bernd Möbius |
INTERSPEECH | 2 |
| 2007 | Tagging syllable boundaries with joint n-gram modelsabstractThis paper presents a statistical method for the segmentation of words into syllables which is based on a joint n-gram model. Our system assigns syllable boundaries to phonetically transcribed words. The syllabification task was formulated as a tagging task. The syllable tagger was trained on syllable-annotated phone sequences. In an evaluation using ten-fold cross-validation, the system correctly predicted the syllabification of German words with an accuracy by word of 99.85%, which clearly exceeds results previously reported in the literature. The best performance was observed for a context size of five preceding phones. A detailed qualitative error analysis suggests that a further reduction of the error rate by up to 90 % is possible by eliminating inconsistencies in the training database. Helmut Schmid, Bernd Möbius, Julia Weidenkaff |
INTERSPEECH | 2 |
| 2007 | Word stress correlates in spontaneous child-directed speech in GermanabstractIn this paper we focus on the use of acoustic as well as voice quality parameters to mark word stress in German. Our aim was to identify the speech parameters parents use to indicate word stress differences to their children. Therefore, mothers and their children were recorded during a period of at least one year while they performed a special playing task using word pairs that differ only in the position of word stress. The recorded target words were analyzed acoustically and with respect to voice quality. The results presented here concern the mothers’ productions of contrastive word stress, and we discuss our findings with respect to the results of previous studies investigating word stress. Our results provide further insight into the process of word stress acquisition in German. Index Terms: production, prosody, word stress, acoustics, voice quality Katrin Schneider, Bernd Möbius |
INTERSPEECH | 2 |
| 2007 | Speaking rate effects in a landmark-based phonetic exemplar modelabstractIn this study we describe a model of speech perception in which neither speaking rate nor lower level temporal cues are considered explicitly. Instead, newly encountered speech signals are encoded as sequences of detailed acoustic events specified in real time at salient landmarks and compared directly with previously heard patterns. When presented with obstruent-vowel sequences occurring in the TIMIT database, the model performs similarly to humans in relying on temporal information for consonant and vowel recognition— and interpreting this information in a rate-dependent manner—when non-temporal cues are ambiguous; and by being adversely affected by local rate variability. These results indicate that compensation for speaking rate in human perception may follow implicitly from even modest knowledge of the robust correlations between temporal and other properties of individual speech events and those of their surrounding contexts, and do not require special normalization processes. Travis Wade, Bernd Möbius |
INTERSPEECH | 2 |
| 2006 | Towards a comprehensive investigation of factors relevant to peak alignment using a unit selection corpusabstractThis paper aims to demonstrate the use of a unit selection corpus, the IMS German Festival synthesis system [1], in carrying out a comprehensive investigation of factors influencing specific aspects of the phonetic realization of tonal categories. The study restricts itself to the alignment of peaks in H*L pitch accents in German. First results confirm not only well-known effects of syllable structure, e.g., peaks occurring relatively early when there is a sonorant onset or relatively late when there is a sonorant in the coda, but also attest to the special status of the nuclear pitch accent vs. accents occurring earlier in the intonation phrase. Furthermore, instances of H*L in syllables directly at the phrase boundaries (initial or final) are shown to behave significantly differently from those that are located farther away. A similar effect is observed when another pitch accent follows the H*L peak in the very next syllable as opposed to a distance of two or more syllables. In these cases it also matters whether a low or high target is following (the peaks occur relatively later when followed by a L target). The results should have the benefit of both describing the specific characteristics of the voice providing the corpus (allowing a more detailed phonetic realization of tonal categories during the synthesis process) and offering general insights into which factors are relevant to the alignment of H*L peaks in German. Index Terms: intonation synthesis, peak alignment, German 1. Matthias Jilka, Bernd Möbius |
INTERSPEECH | 2 |
| 2005 | Perceptual magnet effect in German boundary tonesabstractThe experiment described in this paper tests for the perceptual magnet effect within the categories of high and low boundary tones in German, referring to question and statement, respectively. The experiment is based on previous work in which the categorical status of the two German boundary tones had been evaluated. The results found there showed that there was a discrimination ability within categories which could not be explained by the classical definition of categorical perception. The results reported in the present paper show that a perceptual magnet exists in the statement category but not in the question category. 1. Katrin Schneider, Bernd Möbius |
INTERSPEECH | 2 |
| 2005 | Formant Tracking Using Context-Dependent Phonemic InformationabstractA new formant-tracking algorithm using phoneme information is proposed. Conventional formant-tracking algorithms obtain formant tracks by analyzing the acoustic speech signal using continuity constraints without any additional information. The formant-tracking error rate of the conventional methods is reportedly in the range of 10%-20%. In this paper, we show that if text or phoneme transcription of speech utterances is available, the error rate can be significantly reduced. The basic idea behind this approach is that given the phoneme identity, formant-tracking algorithms can have a better clue of where to look for formants. The algorithm consists of three phases: 1) analysis, 2) segmentation and alignment, and 3) formant tracking by the Viterbi searching algorithm. In the analysis phase, formant candidates are obtained for each analysis frame by solving the linear prediction polynomial. In the segmentation and alignment phase, the text corresponding to the input speech utterance is converted into a sequence of phoneme symbols. Then, the phoneme sequence is time aligned with the speech utterance. A hidden Markov model (HMM) based automatic segmentation algorithm is used for forced-time alignment. For each phoneme segment, nominal formant frequencies are assigned at the center of each phoneme segment. Then nominal formant tracks for the entire utterance are obtained by interpolating the nominal formant frequencies. In order to compensate for the coarticulation effect, different interpolation methods are used depending on the phonemic context. The interpolation process makes the formant-tracking algorithm robust to possible segmentation errors made by the HMM-based segmentation algorithm. As a result, the proposed formant-tracking algorithm does not require highly accurate alignment/segmentation. Finally, a set of formants is chosen from the formant candidates in such a way that the resulting formant tracks come close to the nominal formant tracks while satisfying the continuity constraints. The algorithm is tested using natural speech utterances and the performance is compared against formant tracks obtained by the conventional method using continuity constraints only. The new algorithm significantly reduces the formant-tracking error rate (5.03% for male and 3.73% for female) over the conventional formant-tracking algorithm (13.00% for male and 15.82% for female). Jan P. H. van Santen, Bernd Möbius, Joseph P. Olive |
IEEE Trans. Speech Audio Process. | 3 |
| 2003 | ISCA special session: hot topics in speech synthesis
Gérard Bailly, Nick Campbell 0001, Bernd Möbius |
INTERSPEECH | 3 |
| 2003 | Restricted unlimited domain synthesis
Antje Schweitzer, Norbert Braunschweiler, Tanja Klankert, Bernd Möbius, Bettina Säuberlich |
INTERSPEECH | 4 |
| 2001 | Prosodic models, automatic speech understanding, and speech synthesis: towards the common groundabstractAutomatic speech understanding and speech synthesis, two of the major speech processing applications, impose strikingly different constraints and requirements on prosodic models.The prevalent models of prosody and intonation fail to offer a unified solution to these conflicting constraints.As a consequence, prosodic models have been applied only occasionally in end-toend automatic speech understanding systems; in contrast, they have been applied extensively in speech synthesis systems.In this paper we want to discuss the reasons for this state of affairs as well as possible strategies to overcome the shortcomings of the use of prosodic modelling in automatic speech processing. Anton Batliner, Bernd Möbius, Gregor Möhler, Antje Schweitzer, Elmar Nöth |
INTERSPEECH | 2 |
| 2001 | The ISCA special interest group on speech synthesisabstractThis paper describes the constitution and activities of the ISCA Speech Synthesis Special Interest Group, SynSIG. It summarises past achievements and suggests ways in which future development could be maintained. The aims of the Special Interest Group on Speech Synthesis are to promote the study and diffusion of knowledge about speech synthesis in general, in a number of ways including: dedicated web pages, a mailing list, a bibliographic database, organisation of workshops on specific themes, exchange of students, and helping to co-ordinate sessions on speech synthesis in international conferences and workshops. The international and multi-disciplinary nature of the SIG also provides a means for diffusing information both to and from the different research communities involved in the synthesis of various languages. 1. Nick Campbell 0001, Wolfgang Hess, Bernd Möbius, Jan P. H. van Santen |
INTERSPEECH | 3 |
| 2001 | Towards a model of target oriented production of prosodyabstractA new paradigm for prosody research is presented, inspired by the speech production model recently proposed by Guenther, Perkell, and colleagues. This research paradigm aims at generalizing the production model by extending it from a predominantly segmental perspective to a new theory of the production of prosody. Speech movements in the prosodic domain are interpreted as intonational gestures that are planned to reach and traverse perceptual target regions. Evidence from F0 alignment studies suggests that the perceptual targets can be approximately represented by regions in a multidimensional acoustictemporal space. These studies also indicate that segmental, spectral, temporal, and prosodic structure are co-produced in such a way as to mutually support and enhance, and not impair, the perceptual targets. Furthermore, examples of multilevel mappings between invariant and variable targets in the domain of prosody are provided, and a dichotomy of phonemic and postural prosodic settings is discussed. Grzegorz Dogil, Bernd Möbius |
INTERSPEECH | 2 |
| 2001 | Developments and paradigms in intonation research
Antonis Botinis, Björn Granström, Bernd Möbius |
Speech Commun. | 3 |
| 2000 | Inducing Probabilistic Syllable Classes Using Multivariate ClusteringabstractAn approach to automatic detection of syllable structure is presented. We demonstrate a novel application of EM-based clustering to multivariate data, exemplified by the induction of 3- and 5-dimensional probabilistic syllable classes. The qualitative evaluation shows that the method yields phonologically meaningful syllable classes. We then propose a novel approach to grapheme-to-phoneme conversion and show that syllable structure represents valuable information for pronunciation systems. Karin Müller 0001, Bernd Möbius, Detlef Prescher |
ACL | 2 |
| 1999 | Formant tracking using segmental phonemic informationabstractA new formant tracking algorithm using phoneme dependent nominal formant values is tested. The algorithm consists of three phases: (1) analysis, (2) segmentation, and (3) formant tracking. In the analysis phase, formant candidates are obtained by solving for the roots of the linear prediction polynomial. In the segmentation phase, the input text is converted into a sequence of phonemic symbols. Then the sequence is time aligned with the speech utterance. Finally, a set of formant candidates that are close to the nominal formant estimates while satisfying the continuity constraints are chosen. The new algorithm significantly reduces the formant tracking error rate (3.62%) over a formant tracking algorithmusing only continuity constraints (13.04%). We will also discuss how to further reduce the tracking error rate. INTRODUCTION In the Bell Labs' Text-To-Speech (TTS) system [1], a limited number of acoustic units is stored in the inventory table. Therefore, it is important to be able to... Jan P. H. van Santen, Bernd Möbius, Joseph P. Olive |
EUROSPEECH | 3 |
| 1999 | The Bell Labs German text-to-speech system
Bernd Möbius |
Comput. Speech Lang. | 1 |
| 1998 | Contextual effects on voicing profiles of German and Mandarin consonantsabstractIn this paper we present a study of the voicing profiles of consonants in Mandarin Chinese and German. The voicing profile is defined as the frame-by-frame voicing status of a speech sound in continuous speech. We are particularly interested in discrepancies between the phonological voicing status of a speech sound and its actual phonetic realization in connected speech. We further examine the contextual factors that cause voicing variations and test the cross-language validity of these factors. The result can be used to improve speech synthesis, and to refine phone models to enhance the performance of automatic speech segmentation and recognition. Chilin Shih, Bernd Möbius |
ICSLP | 2 |
| 1997 | The bell labs German text-to-speech system: an overviewabstractIn this paper we present an overview of the German version of the Bell Labs text-to-speech system, a high-quality concatenative synthesis system with extensive text analysis capabilities. We discuss problems of text analysis, and our solutions to these problems, including: the integration of text normalization tasks into linguistic text analysis; the capability to morphologically analyze compounds and unseen words; name analysis and pronunciation. We briefly describe the prosodic components of the text-to-speech system and their underlying duration and intonation models. Finally, the phonetically motivated structure of the acoustic inventory is presented. Bernd Möbius, Richard Sproat, Jan P. H. van Santen, Joseph P. Olive |
EUROSPEECH | 1 |
| 1997 | Multi-lingual duration modelingabstractControlling timing in text-to-speech synthesis systems is complicated, because there are many contextual factors that affect timing; moreover, which factors matter and what their precise effects are varies among languages. We describe here a language-independent approach for duration control. At run time, a language-independent timing module accesses languagespecific tables. These tables specify which sub-classes of the feature space (i.e., all combinations of context and phone identity) are homogeneous in the specific sense that the same factors have similar effects on the cases in a sub-class. Within a sub-class, durations are modeled by simple arithmetic models such as multiplicative, additive, or – more generally – sums-ofproducts models. Exploratory statistical methods (supervised) and parameter estimation techniques (unsupervised) are used for Jan P. H. van Santen, Chilin Shih, Bernd Möbius, Evelyne Tzoukermann, Michael Tanenblatt |
EUROSPEECH | 3 |
| 1996 | Modeling segmental duration in German text-to-speech synthesis
Bernd Möbius, Jan P. H. van Santen |
ICSLP | 1 |
| 1993 | Analysis and synthesis of German F0 contours by means of Fujisaki's model
Bernd Möbius, Matthias Pätzold 0002, Wolfgang Hess |
Speech Commun. | 1 |
| 1992 | F0 synthesis based on a quantitative model of German intonation
Bernd Möbius, Matthias Pätzold 0002 |
ICSLP | 1 |